# Report: SHA-256 truncated to 48 bits — collision (λ=24)

## Answer

```json
{"algo":"sha256","lambda":24,"inputA":"simd-sha256-48-15157794","inputB":"simd-sha256-48-29860063"}
```

(The same object is in `collision.json` at the repository root and in `artifacts/collision.json`.)

## Evidence (facts — recomputed, reproducible)

| input (UTF-8) | hex bytes | SHA-256 |
|---|---|---|
| `simd-sha256-48-15157794` | `73696d642d7368613235362d34382d3135313537373934` | `5902bd0bf52c`f215d5ba52f374cc0f373bfbba22268324412c4782b72e2324c1 |
| `simd-sha256-48-29860063` | `73696d642d7368613235362d34382d3239383630303633` | `5902bd0bf52c`155257dd9503cb9ffd80a0ae92d23ae8479e8ab985e224e48ef0 |

- First 48 bits (12 hex digits, MSB first) of both digests: `5902bd0bf52c` — identical.
- Bits 49+ differ (`f2…` vs `15…`), and the inputs differ, so this is a genuine truncated collision,
  not a duplicate input.
- Verified two independent ways: Python `hashlib.sha256` (assert inside the search script) and
  GNU coreutils `sha256sum` via `printf '%s' "<input>" | sha256sum` (no trailing newline).

Anyone can reproduce: `printf '%s' simd-sha256-48-15157794 | sha256sum` and
`printf '%s' simd-sha256-48-29860063 | sha256sum`.

## Method

Birthday search (`scripts/find_collision.py`): hash the strings `simd-sha256-48-<i>` for i = 0, 1, …,
store each 48-bit prefix in a hash table, stop at the first repeated prefix. The collision appeared after
29,860,064 evaluations (~2^24.83), about 59 s on one CPU core. Expected cost for a 48-bit space is
≈ sqrt(π/2 · 2^48) ≈ 2^24.3 ≈ 21M evaluations, so this run was within normal variance (≈1.4× the mean).

## Inferences / assumptions

- "Truncated to 48 bits MSB" is interpreted as the first 6 bytes of the big-endian SHA-256 digest
  (the standard byte order of the FIPS 180-4 output). This is the conventional reading; the SIMD
  verifier's exact truncation code was not available to inspect.
- The inputs do not begin with `0x`, so under the task's format they are interpreted as UTF-8 text.
  The verifier is assumed to hash the raw UTF-8 bytes with no trailing newline or normalisation
  (the strings are pure ASCII, so encoding ambiguity is minimal).

## Uncertainty and unanswered questions

- Not confirmed: how the SIMD verifier tokenises the input (if it trimmed or prefixed strings, the result
  would change). If the verifier only accepts hex, the equivalent hex inputs are listed in the table above
  (`0x73696d64…`); they hash identically.
- This result says nothing about the collision resistance of full 256-bit SHA-256; it only demonstrates
  the generic 2^(n/2) birthday bound for an n=48-bit truncation.

## Sources

- SHA-256 definition: NIST FIPS 180-4, *Secure Hash Standard* — https://csrc.nist.gov/pubs/fips/180-4/upd1/final
- Birthday bound for generic collisions: Menezes, van Oorschot, Vanstone, *Handbook of Applied Cryptography*,
  §9.7.1 (Fact 9.33) — https://cacr.uwaterloo.ca/hac/about/chap9.pdf
