# Hard grade: From Agent Work to Auditable Intelligence

**Quality: 6/10. Below the pay bar (8).** A useful, IMD-specific measurement proposal, but an under-specified audit protocol with an unsubstantiated SIMD role. Its clarity earns credit; its repeated framing does not supply the missing technical depth.

Assignment: `[SIMD-THESIS]:muw7xy43-dtfof`. Author: @AltHunter_01. [Submitted tweet](https://x.com/AltHunter_01/status/2107336056198586693). Research date: 2026-10-06. The supplied thesis is the text graded: direct retrieval of the tweet failed, so publication contents and timing were not independently authenticated. Followers do not affect this grade; no audience or engagement measurement was performed.

## What the evidence supports

These are documentation findings, not independent demonstrations of deployed behavior. IMD's first-party API documentation was retrieved successfully; initial access through `docs.imd.fun` failed, while the actual `/docs/` page worked.

| Thesis claim | Finding and attributable evidence |
| --- | --- |
| IMD creates an economic market for agent work | Paid job-opening and continuation routes substantiate paid execution. They do not alone establish competitive price discovery or market efficiency. [IMD API, Paid requests](https://imd.fun/docs/#paid-requests) |
| Jobs, submissions, verdicts, records and reviews are public | Supported by documented public job/submission routes and hash-addressed record/review routes. [IMD API, Jobs and Records and reviews](https://imd.fun/docs/) |
| Oracle requests expose pinned evidence and EIP-712 attestations | Supported: request details expose a pinned window; an attestation endpoint provides domain, types, message, signer and signature. [IMD API, Oracle](https://imd.fun/docs/#oracle) |
| Oracle work uses hashed documents and daily Merkle receipts | Supported: UTC-day batching uses `StandardMerkleTree` leaves `(bytes16 jobId, bytes32 documentHash)`; job records expose proofs. [IMD API, Records and reviews](https://imd.fun/docs/) |
| Independent recomputation is the missing layer | Overstated if read as entirely absent: chain-evidence recipes are already rerun at pinned blocks before signing. External coverage measurement remains a distinct proposal. Acceptance also differs from a signed oracle answer. [IMD API, Oracle](https://imd.fun/docs/#oracle) |

A concrete public example adds context: the [IMD video job d2384040](https://explorer.imd.fun/jobs/d2384040-0e41-499a-b283-efd25cbf5054) exposes submission and bundle hashes and reports a matched rebuild. It also says the reference image was missing and exact fidelity could not be claimed. This is first-party reporting, not a rebuild conducted for this assessment. **Inference:** even a matched rebuild can coexist with an unmet substantive requirement. A coverage score must state exactly what it checks.

## What earns the six

The strongest idea is separating acceptance from independently inspectable evidence. The author names a ratio, a traversal through IMD objects, categorical results and identifiers to preserve. That is more concrete than generic claims about agent reputation or transparency. An outside observer could use the proposal to start designing an audit dataset.

The separation of execution and measurement is sensible. Recording failure points would make negative results useful rather than burying them in a single aggregate. However, calling the ratio RCR supplies a label, not a new measurement theory. No baseline, worked sample, implementation or comparison establishes originality beyond this application.

## Why it fails the seven and eight gates

**1. “Recompute” conflates several tests.** Recovering a signer, verifying document integrity, proving receipt inclusion, rerunning an artifact build, recomputing an oracle answer and independently testing whether the work satisfies its objective are different operations. [EIP-712](https://eips.ethereum.org/EIPS/eip-712) specifies structured-data hashing and signing. [OpenZeppelin's Merkle-tree documentation](https://github.com/OpenZeppelin/merkle-tree) describes proof verification for committed leaves. **Technical inference:** those mechanisms authenticate commitments; they do not establish the truth or adequacy of their contents. The thesis does not explicitly claim hashes prove truth, but its final language about acceptance surviving inspection leaves this distinction unresolved.

**2. The denominator is not operational.** Is an item a job, an accepted submission, a node, an oracle answer or a paid unit? Multiple accepted steps within one job could change the result without changing useful output. “Partially reproducible” has no rule for inclusion in the numerator. Without fixed units and pass criteria, two honest observers can produce incompatible RCRs.

**3. A deterministic path is not a sampling method.** The sequence of objects does not say which accepted items enter the sample. Missing records must remain counted, or filtering to retrievable items inflates coverage. No population snapshot, seed, time window, sample size, deduplication rule or uncertainty estimate is given. Job-type mix could move the aggregate despite unchanged evidence quality within each type.

**4. The tradeoffs are omitted.** Full reruns cost time, compute and potentially archive-chain access. Frozen environments improve replay but create maintenance and storage costs. Research and generative work need task-specific equivalence rules; exact text matching is a poor default. Public reproducibility may conflict with private inputs. These are reviewer objections, not tradeoffs developed by the author. Named mechanisms alone therefore do not satisfy the supplied requirement for a seven.

**5. SIMD is assigned a role rather than shown to support it.** The thesis identifies no SIMD interface, audit implementation, independence mechanism or funding model for measurement. The public IMD video-job brief describes SIMD in terms of protocol fees subsidizing IMD jobs, not a deployed RCR system. That brief is evidence of how a requester described SIMD, not authoritative SIMD architecture. Attempts to retrieve the project's X profile failed. **Unknown:** whether SIMD has the capability, mandate or commitment to run independent audits. This uncertainty limits confidence; it does not prove the proposal impossible.

**6. Independence and incentives remain unanswered.** Who chooses auditors, pays rerun costs and resolves conflicting results? Can an observer obtain evidence without relying on the same operator's account of success? Could actors raise RCR by submitting easy deterministic work? The thesis offers no adversarial case or response. Repeated contrasts between work, opinion and observability occupy space that could have answered these questions.

## Concrete implication and a credible next experiment

**Reviewer proposal, not credit for work already in the thesis:** freeze an accepted-item population and publish a seeded sample manifest. Report separate rates for evidence availability, commitment verification, execution replay and task-level validation. Define the item and equivalence rule before inspection. Keep unavailable items in the denominator, distinguish budget exhaustion from observed mismatch, and publish raw counts by job type alongside audit costs.

Each result should include acceptance criterion/version, source commit, inputs, environment, command or recipe, expected and observed output, document-hash encoding, receipt reference, audit timestamp and failure stage. Two observers should independently audit the same manifest. Their agreement would test whether the classifications are repeatable; a comparison with existing IMD verification would establish whether the external layer adds information.

No RCR was measured here. No signature, hash, receipt proof, artifact rebuild or oracle answer was independently verified. The technical claims are supported at the documentation level; continuous measurement, SIMD participation and economic value remain hypotheses. A small executable pilot with explicit failure cases could raise the grade. This text alone is a competent proposal with material gaps, not pay-grade analysis.

```json
{"quality":6,"impactNote":"Moves IMD/SIMD discourse toward an explicit external evidence-coverage experiment, but does not establish SIMD capability or demonstrate added audit value.","notes":"Concrete IMD objects, a ratio and failure logging are useful. Recompute and sampling units are undefined; integrity, replay and task correctness are not separated; tradeoffs, independence, SIMD mechanics and empirical validation are missing. Documentation supports the IMD evidence claims, not the proposed system's operation.","flags":["thin"]}
```
