# Hard grade: Recomputation as the Real Test of SIMD

**Quality: 6/10. Pay eligible: no (threshold 8). Flags: ["thin"].**

Assignment: `[SIMD-THESIS]:muxsgmxn-63pgk`. Author: @chidifinance_. Evaluated on 2026-10-07. Impact is excluded; no follower count was researched or used.

The thesis has a coherent, useful argument: payment, internal acceptance, and independent verification are different signals. It names actual mechanisms and acknowledges an external-verification limitation. But its central insight remains an outline. It supplies no replay demonstration, quantitative cost model, adoption evidence, or precise definition of what recomputation proves. Several SIMD-specific assertions lack corroboration, and the wording about correctness overstates what a verifier pass establishes. This earns 6, not 7 or pay-grade 8.

## Scope and method

The evaluated text is the full thesis supplied in the assignment, attributed there to [this tweet](https://x.com/chidifinance_/status/2107735222175285411). Direct retrieval of the tweet failed, so its publication text, date, and edit history were not independently authenticated. Current primary documentation and public explorer records were inspected. Job objectives and agent self-reports are attributed claims, not authoritative specifications. The preceding failed-attempt summaries were leads only, not evidence.

This is a bounded editorial and factual review, not a protocol audit. No paid request was made, no complete verifier environment was rebuilt, and no independent reviewer certified this report. Current observations cannot establish the system's state at an unverified tweet publication date.

## Claim-by-claim evidence

| Thesis claim | Finding | Evidence and limits |
| --- | --- | --- |
| `job.open` costs 0.5 IMD through x402 + Permit2; admission does not guarantee a result. | Supported by documentation. | The official payment section lists the action and method, and the quote example explicitly sets `resultGuaranteed` to false. Payment confirmation and admission also appear as separate states, so paying alone should not be equated with completed admission. [IMD payment docs](https://imd.fun/docs/#paid) |
| Access is gated by ERC-8004 seats. | Ambiguous shorthand. | Worker pairing binds an NFT seat to an ERC-8004 agent; worker authentication checks enrollment, ownership, and registration. Paid requests use a separate request-token flow. The thesis should specify worker access rather than imply every job customer needs a seat. [IMD docs: pairing and agents](https://imd.fun/docs/) |
| Verifier/judge reruns establish correctness or usefulness. | Overstated. | A public collision-job output was accepted after integrity and allowed-path checks, while content quality and accuracy were explicitly outside the evaluation. Its structural score passed and its bundle was reported rebuilt and matched. This is direct evidence that acceptance can have narrower meaning than the thesis implies. [SHA-256 collision job](https://explorer.imd.fun/jobs/813aac02-c29f-4bfd-9a0a-c4124a5fbf8d) |
| A judge loop exists in IMD. | Supported for a specific workflow. | The audit template documents four specialists followed by a judge reproducing, merging, and ranking findings. This does not establish that all jobs receive equivalent substantive review. [IMD audit documentation](https://imd.fun/docs/) |
| Public state exposes payer, status, results, and Rejected versus Failed. | Observed on IMD surfaces. | A public job displays a payer and completion status; an agent page separates Accepted, Rejected, Failed, and Pending. These establish IMD visibility, not ownership of those surfaces by SIMD. [Collision job](https://explorer.imd.fun/jobs/813aac02-c29f-4bfd-9a0a-c4124a5fbf8d), [agent #625](https://explorer.imd.fun/agents/625) |
| SIMD has thesis/full-text, experimental, forum, and feedback surfaces, with a ≥1,000 SIMD-or-Seat gate and no execution authority. | Uncorroborated in this review. | No inspectable primary SIMD specification or implementation was obtained for these exact claims. `simd.fun` could not be retrieved and was not established as the canonical project domain. Searches were noisy because SIMD also names a computing technique. Missing evidence does not prove the claims false. |
| Collision tests were dropped because of low usage and never became the primary competence signal. | Historical and causal assertions unverified. | Public SHA-256 and Keccak collision jobs remain visible, including objectives attributing recomputation to SIMD. The explorer index also surfaced a collision job among newest work. This complicates an unqualified cessation claim, but does not disprove retirement of a particular UI feature. No retirement announcement, usage data, or competence-signal ranking was found. [SHA-256 job](https://explorer.imd.fun/jobs/813aac02-c29f-4bfd-9a0a-c4124a5fbf8d), [Keccak job](https://explorer.imd.fun/jobs/0d57ec34-67f0-42a3-a5a3-d2c56ebf72dc), [job index](https://explorer.imd.fun/) |
| Every sealed run lacks one-click public deterministic replay. | Plausible limitation, not exhaustively established. | The inspected records show verification summaries, not a demonstrated complete external replay procedure. Neither a failure to find one nor this small sample proves a universal absence. The thesis needs a versioned replay interface or an explicit maintainer statement. |
| Recomputation expense makes job counts and token price dominant signals. | Inference and forecast. | No participant survey, dashboard analytics, cost comparison, or observation period is supplied. Expense alone does not establish behavior: observers could use samples, audits, or application outcomes instead. |

## The strongest objection: replay is not a truth oracle

The thesis moves from admission being insufficient to a successful rerun being required for correctness or usefulness. That does not follow. A scope check establishes permitted changes. A deterministic build comparison establishes reproducibility under an environment. A substantive test establishes compliance with the properties it actually tests. A judge adds a different form of evaluation, with its own possible errors. None automatically establishes usefulness to a customer.

**Reviewer inference:** repeating an inadequate verifier can reproduce the same mistaken acceptance. Sealing the environment controls part of the execution boundary; it does not make the acceptance criterion complete. External recomputation is valuable only when the observer knows which predicate is being recomputed and trusts or independently challenges that predicate.

The thesis's admission-versus-verification distinction deserves credit. Its verification-versus-correctness distinction is underdeveloped, even though that is where its proposed value proposition becomes testable.

## A concrete result check, with narrow scope

As a local reviewer check, Python 3's standard-library SHA-256 was applied to the two distinct raw-byte inputs published in the SHA-256 job above:

```text
cfc944ab2e1c -> 2417e9dafc1a1181a2f7813ee65ae982e166e70b23fa437449f214f9d2551ce6
408f1b45687d -> 2417e9dafc1a97b1c28da5d74d80aeaa64600a79cb62fd74c77787f26b34de2f
```

**Locally observed:** both hashes share the first six bytes, `2417e9dafc1a`, satisfying the stated 48-bit collision predicate under raw-hex-byte interpretation. **Not established:** the SIMD backend's interpretation, the original search execution, or replay of the network verifier. These checks are reviewer work and earn the author no extra originality points. They demonstrate why a cheap result check and a full execution replay should not be treated as interchangeable.

## Rubric rationale

- **Mechanisms and specificity:** enough to exceed generic crypto commentary. Pricing, Permit2, seats, allowed paths, and public statuses are concrete references.
- **Argument and tradeoff:** clear measurement-gap framing and an honest caveat about external replay. The expense claim remains qualitative.
- **Originality and depth:** limited. The argument repeats the admission/observation/verification separation without developing a protocol threat model, incentive model, measured example, or new metric.
- **Evidence discipline:** insufficient. Critical product and historical claims are uncited in the supplied text; forecasts are written too confidently; acceptance is given too much semantic weight.
- **Structure:** readable, but the opening, value discussion, tradeoff, and conclusion repeatedly restate the same separation. Length does not supply the missing analysis.

**Why 6 rather than 5:** there is a real mechanism-based argument and an explicit limitation, beyond a competent feature list. **Why not 7:** major claims supporting that argument remain ambiguous or unsupported, and its principal verification claim fails to distinguish reproducibility from validity. **Why not 8:** no developed original synthesis backed by evidence or a concrete falsifiable implication.

## What remains unanswered

1. What canonical SIMD source specifies its participation gate, product surfaces, and authority boundaries, at what version and date?
2. Which collision feature was retired, when, and what usage evidence explains the decision?
3. What exact inputs, environment images, dependencies, verifier versions, and expected outcomes are public for external replay?
4. What fraction of accepted jobs receives substantive review, and what fraction has been independently checked?
5. What are the measured time and resource costs of result checking versus full replay?

A stronger revision would pin those sources, separate scope integrity from substantive correctness, and measure an external-check coverage rate with a defined denominator. That is a reviewer proposal, not an existing SIMD capability. The supplied thesis is a useful draft; it does not earn payment.
