# Report: keccak256 truncated-48-bit collision (λ=24)

## Answer
```json
{"algo":"keccak256","lambda":24,"inputA":"imd-9804673","inputB":"imd-17368822"}
```

| input (UTF-8) | keccak256 digest |
|---|---|
| `imd-9804673`  | `2f72fe059bc2`493e489ac23cd5bb0119eee739455c6ba49cfd31f75a2b7be703 |
| `imd-17368822` | `2f72fe059bc2`f734345293e1223f429e97aea53934f75f186ac1c57ca5f1d550 |

The first 6 bytes (48 bits) are the same: `2f72fe059bc2`. The inputs are different strings.

## Evidence (facts, reproduced locally)
1. **Search:** `tools/collide.c` hashed `"imd-0"`, `"imd-1"`, … with Keccak-256 and stored each 48-bit prefix in an open-addressing hash table. It reported a match after **17,368,823 evaluations** (≈2^24.05) in about 96 s on one core. Output: `A=imd-9804673 B=imd-17368822 prefix=2f72fe059bc2`.
2. **Self-test of the C implementation:** it reproduces the published Keccak-256 vectors: `""` → `c5d24601…5d85a470` and `"abc"` → `4e03657a…a12d6c45`.
3. **Independent re-check:** `tools/verify.py` is a separate pure-Python Keccak (different code structure). Checks it passes:
   - Its permutation with SHA3 padding (0x06) matches Python's `hashlib.sha3_256` on 3 inputs, including a multi-block one.
   - Its Keccak-256 output matches the two vectors above.
   - On `collision.json` it recomputes both digests (shown in the table) and prints `48-bit prefix collision: True`.

## Inferences
- The evaluation count fits the birthday bound. The expected count for a 48-bit space is about √(π/2 · 2^48) ≈ 2.1·10^7, and we used 1.74·10^7. This is a sanity check, not proof. The proof is that the digests can be recomputed.
- "keccak256" was read as Ethereum-style Keccak-256 (padding 0x01) because that is what the name usually means. Under this reading the empty-string vector `c5d2…a470` is correct.

## Uncertainty / open questions
- **Hash variant:** if the SIMD verifier actually uses FIPS-202 SHA3-256 under the name "keccak256", this pair would almost certainly **not** collide. I did not test it as a SHA3-256 collision.
- **Truncation rule:** "first 48 bits MSB" was taken to mean the first 6 bytes of the digest in its standard byte order. If the verifier truncates differently, the result may not hold.
- **Input encoding:** the inputs have no `0x` prefix, so they should be read as UTF-8. I could not confirm how the verifier parses inputs beyond the task text.
- No external sources were needed. All claims above can be reproduced with the two scripts in `tools/`.
