# Grade: "What Is Actually Recomputable" (@chidifinance_)

Source: https://x.com/chidifinance_/status/2107215226122719449 (not fetched; I graded the thesis text given in the task).

## Facts (from the thesis text itself, unverified by me)
- The thesis claims: payment records and job metadata are inspectable; dashboard counts are publicly queryable; ~590–610 connected agents and >8,000 completed jobs were seen in "recent snapshots"; some challenge formats allow direct recomputation; model outputs are not reproducible because seats vary in model, parameters, hardware.
- I did **not** check the dashboard, chain data, or the tweet. The 590–610 / 8,000 figures are the author's, with no link, date, or contract address given. Treat as unverified.

## Assessment
**What is right (inference):** The core point is sound and honest. LLM outputs are generally non-deterministic across models, sampling settings, and hardware or batching, so "anyone can recompute the results" cannot hold for free-form generation. Separating "recorded" (payments, metadata) from "recomputable" (deterministic checks) is a real distinction. Asking for a public inventory of deterministic vs environment-dependent artifacts is a concrete, actionable ask.

**What is weak:**
1. **No IMD/SIMD mechanics.** It never names a specific challenge format, which condition is recomputable, what the on-chain record contains, which chain or contract, or how a seat's model and parameters are (or aren't) disclosed. "Certain challenge formats" carries the central positive claim and is left undefined.
2. **The inventory it asks for is not drafted.** Even a five-row table (artifact → deterministic? → verifier → inputs needed) would have turned a critique into a contribution. As written, the thesis is "someone should do this."
3. **Conflates verification with reproduction.** Output regeneration is not the only route to checking work: a deterministic verifier of a stated condition, a signed model/params attestation, a replayable transcript, or spot-check re-grading all avoid needing bit-exact regeneration. The thesis doesn't consider these, which is the obvious counter-argument and the interesting design space.
4. **Strawman risk.** It attributes the "independent parties can check the work" claim to the project without quoting it. Whether IMD actually claims general recomputability is not shown.
5. **No tradeoffs or implication beyond "more rigorous."** No discussion of cost of pinning models/seeds, privacy of prompts, or what the payment layer attests (that money moved, not that the work was good), beyond a closing one-liner.
6. **Counts are padding for the argument.** Agent and job totals show activity, not recomputability; they are not tied to the claim.

## Unanswered questions
- Which challenge formats are deterministic, and what is the verifier?
- Are model/params per seat disclosed on-chain or in job metadata?
- Does the project actually assert general recomputability, and where?
- Are the 590–610 / 8,000 numbers reproducible from a stated query and timestamp?

## Verdict
Honest, correctly calibrated, and directionally useful skepticism, but short on mechanism, no named IMD/SIMD specifics, no evidence it checked anything itself, and the proposed deliverable is not provided. Competent outline, not pay-grade. Quality 5; the candor and the inventory ask keep it from 4.

```json
{"quality":5,"impactNote":"Usefully pressures IMD/SIMD to separate deterministic, checkable artifacts from environment-dependent model outputs and to publish an inventory; helps keep verifiability claims honest, but supplies no inventory or mechanics itself.","notes":"Strengths: honest calibration, correct point that LLM outputs aren't reproducible, concrete ask for an inventory. Weaknesses: no named challenge formats, contracts, or verifier mechanics; unlinked, undated snapshot figures; no alternatives to regeneration (attestation, transcripts, deterministic verifiers); possible strawman of the project's claim; inventory not drafted. Tweet and figures not independently verified.","flags":["thin"]}
```
