# SIMD-COLLISION sha256 / λ=24: 48-bit truncated collision

## Answer

```json
{"algo":"sha256","lambda":24,"inputA":"simd-3463940","inputB":"simd-6587109"}
```

This is the exact content of `collision.json` in the repository root. Both inputs are
UTF-8 strings (12 ASCII bytes each, no trailing newline, no `0x` prefix).

## Evidence (facts, reproduced locally)

| Input (UTF-8) | Bytes as hex | SHA-256 |
|---|---|---|
| `simd-3463940` | `0x73696d642d33343633393430` | `3441ac04295a` `c59936fe7b0f52aa8c241ecee4a588e2c4298b3f93217ad81a55` |
| `simd-6587109` | `0x73696d642d36353837313039` | `3441ac04295a` `6cc872775528a28dffd1e26de373387d7b80de925476470b7383` |

- The inputs are distinct (they differ in 7 characters).
- The first 48 bits (6 bytes, most significant first) of both digests are `3441ac04295a`.
- The digests differ from byte 7 onward, so this is a truncated collision only, not a
  full SHA-256 collision.
- The digests were computed two independent ways: Python `hashlib.sha256` inside the
  search script (which asserts equality before writing the file) and coreutils
  `sha256sum` on `printf %s <input>`. Both agree.

Re-check without the search:

```sh
printf %s simd-3463940 | sha256sum
printf %s simd-6587109 | sha256sum
```

## Method

`tools/find_collision.py` (Python standard library only) hashes the 2^25 strings
`simd-0` … `simd-33554431`, keeps the 6-byte digest prefix of each, and looks for repeated
prefixes. The run took 38 s and found 5 colliding pairs; the one with the smallest first
index was written out. The search is deterministic, so rerunning it reproduces the same
file.

All five pairs found (indices `n` in `simd-<n>`): (3463940, 6587109),
(12903921, 30780106), (16041887, 33346017), (16821227, 24798892), (19414987, 23000629).
Only the first was re-verified with `sha256sum`; the others were detected by the script
but not checked separately.

## Inferences

- Finding 5 pairs is in line with the birthday estimate: 2^25 inputs give about
  2^49 pairs, each colliding on 48 bits with probability 2^-48, so about 2 are expected.
  This is a consistency check, not a statistical test.

## Uncertainty and open questions

- **Input encoding as seen by the SIMD verifier.** The task allows "hex 0x... or utf8".
  I chose plain UTF-8 strings without a `0x` prefix and assume the verifier hashes the
  UTF-8 bytes of such a string exactly as written. I could not run the SIMD verifier, so
  this is unconfirmed. If it requires hex, the equivalent inputs are the `0x…` values in
  the table above, which denote the same bytes.
- **"48 bits MSB"** is read as the first 6 bytes of the digest in its standard big-endian
  byte order. Under any other reading (for example the last 6 bytes) this pair does not
  collide.
- **File location.** The task names the file `collision.json` but gives no directory; it
  is committed at the repository root.
