# SHA-256 truncated to 50 bits: collision report (λ=25)

## Answer

```json
{"algo":"sha256","lambda":25,"inputA":"imd-14354348","inputB":"imd-48319665"}
```

Both inputs are plain UTF-8/ASCII strings (no `0x` prefix, so they are not hex-decoded).
The same JSON is in `artifacts/collision.json` and `collision.json` at the repository root.

## Evidence (facts, recomputed locally)

Full SHA-256 digests (`printf '%s' <input> | sha256sum`):

| input | SHA-256 |
|---|---|
| `imd-14354348` | `1bf95fda3510a8afbe8eb0b908153e8a76d59609da8e03452d301e7173df0d3d` |
| `imd-48319665` | `1bf95fda3510b949aa301f837200e27ca475e9afa68c7336855653ef15d1017c` |

The first 50 bits are the first 12 hex digits (48 bits) plus the top 2 bits of the 13th digit.
`1bf95fda3510` is the same in both. The 13th digits are `a` = `1010` and `b` = `1011`, so both
start with `10`. The bits diverge only at bit 51. As an integer, digest >> 206 = `0x6fe57f68d442`
for both inputs. A Python `hashlib` check confirmed this, and it also confirmed that the inputs differ.

Two more collisions came out of the same run. They were verified the same way and are not used in
the answer:
- `imd-5216262` / `imd-31351868` → `0x856a4c61e0be`
- `imd-31224531` / `imd-41039209` → `0x2f2e31122c916`

## Method

A generic birthday search; no weakness of SHA-256 is used.
1. I hashed 2^26 inputs `imd-0` … `imd-67108863` with Python `hashlib.sha256` and kept the top
   50 bits of each digest (about 2 min 13 s).
2. I sorted the (prefix, index) pairs in C with `qsort` and listed the adjacent equal prefixes
   (about 31 s).
3. With N = 2^26 samples in a space of 2^50, the expected number of colliding pairs is
   N²/2^51 = 2. The run found 3, which is consistent with that.

## Inference and uncertainty

- **Assumption:** the verifier takes "truncated to 50 bits MSB" to mean the first 50 bits of the
  big-endian digest, as in standard SHA-256 output order. If it truncated from the least-significant
  end instead, this pair would not collide. That would be an unusual reading of "MSB".
- **Assumption:** inputs without a `0x` prefix are hashed as their UTF-8 bytes, with no trailing
  newline. I chose ASCII-only strings so that no encoding ambiguity can arise.
- **Open question:** I have not seen how the external "SIMD" verifier is implemented. The local
  checks above are the only evidence, and they carry no independent authority.

## Sources

There are no web sources. All claims are reproducible computations: anyone can check them with
`sha256sum` or any SHA-256 library on the two strings above. The SHA-256 definition is NIST FIPS 180-4
(https://csrc.nist.gov/pubs/fips/180-4/upd1/final).
