# Hard grade: "Measure the marginal value of adding agents" (@chidifinance_)

- **Tweet:** https://x.com/chidifinance_/status/2107602310083817709 (could not be fetched: HTTP 402. Graded from the thesis text supplied with the task.)
- **Author:** @chidifinance_ (follower count from the task, ≈307. Not checked, not used in the score.)
- **Graded:** 2026-10-06

## Verdict

**QUALITY: 5 / 10**

The thesis is competent and clearly about IMD/SIMD, and it lists a usable set of metrics. Its central idea, though, is the textbook diminishing-returns / redundancy-versus-coverage tradeoff. The one piece of evidence it offers, the CallBook audit, is unverified. The thesis also misdescribes that evidence: IMD's audit template uses four *differently-scoped* specialists, not four redundant copies of one reviewer. Finally, the prediction is too hedged to be cleanly falsified. That puts it at the top of the "5" band. It does not reach "6" because the one IMD-specific example actually works against the redundancy framing it is used to support.

## What the primary sources show (facts)

| # | Fact | Source |
|---|------|--------|
| F1 | IMD audit jobs run **four specialist nodes with different scopes**: `audit_permissions` (access control), `audit_flow` (control flow and logic), `audit_math` (math and accounting) and `audit_economics` (incentives and token mechanics). | [IMD API docs](https://imd.fun/docs/) |
| F2 | An `audit_judge` node merges, deduplicates, ranks by severity and reproduces findings with proofs. It publishes `GET /jobs/:id/report.md`. Accepted submissions expose findings with location, reproduction and Foundry proof data via `GET /jobs/:id/submissions`. | [IMD API docs](https://imd.fun/docs/) |
| F3 | The public agent explorer shows per-agent status, tasks accepted, success rate, turns, version, hours and wallet. The listing does **not** show per-job findings, specialist counts or judge compression. | [IMD explorer](https://explorer.imd.fun/agents) |
| F4 | The SIMD homepage says only that agents are "followed in real time, as jobs open and move through the swarm". It names no metrics. | [superimdc.xyz](https://www.superimdc.xyz/) |
| F5 | Swarm-wide numbers as of 2026-09-25: about 50,700 attempts, 86% acceptance, an adversarial peer-review stage, and a verifier that rebuilds submissions in a sandbox. | [Bankless](https://www.bankless.com/read/inside-imd-ethereum-s-new-ai-swarm-experiment) |
| F6 | No public source I could find mentions a "CallBook" audit or its 24→8 figures. | Web searches on 2026-10-06 (no matching result) |

## Assessment (inference, tied to the facts above)

### What earns credit
1. **IMD-specific and mechanism-aware.** The thesis correctly describes the specialist → judge → reproduction pipeline (consistent with F2). That is more than generic crypto content.
2. **Concrete metric list.** It proposes recording agents, specialist findings, unique findings, reproduced findings, severity-weighted accepted findings and judge compression ratio. All of these could be derived from data IMD already exposes per job (F2), even though SIMD does not currently show them (F3, F4). This is the most useful part of the post.
3. **A sensible nuance.** It says outright that a high compression ratio is not automatically waste, because convergence feeds validation. This avoids the naive reading that "24→8 means 67% waste".
4. **Defines a measure.** "New, independently supported accepted findings per additional agent after the first" is a workable definition.

### What costs points
1. **The example misreads the architecture (biggest problem).** Per F1, the four CallBook agents were *not* interchangeable reviewers of the same scope. They were permissions, flow, math and economics specialists. Overlap between them is cross-domain convergence: the same root cause seen from different angles. It is not "repeatedly rediscovering the same issue" by redundant labour. The thesis builds its redundancy argument on the one data point that least fits it.
2. **Agent count barely varies where the example lives.** The audit template fixes four specialists (F1). "Compare jobs with different agent counts within comparable task classes" needs natural variation that the audit class may not have. The thesis never asks whether that variation exists or how to create it (for example, controlled ablations that drop one specialist).
3. **24→8 is uninterpretable as stated.** The thesis does not separate (a) duplicates merged, (b) invalid or false-positive findings rejected, and (c) low-severity findings folded together. Those three have opposite implications for "marginal value", yet the post treats compression as one quantity.
4. **The CallBook figures cannot be verified (F6).** There is no job ID, report link or submission link, although IMD publishes exactly these (F2). Claims such as "a high-severity issue independently surfaced and then reproduced" therefore rest on the author's word.
5. **The prediction is weak.** "If adding agents increases accepted findings faster than duplicates → information gain; if it plateaus → redundancy" is close to a restatement of the definitions. There is no threshold, no expected curve shape, no sample size and no control for confounders (codebase size, task difficulty, ground-truth defect count). As written, no data could embarrass it.
6. **Thin originality.** Diminishing returns from extra reviewers, and overlap-based yield estimation (the capture–recapture idea from software inspection research), are well-known concepts. The post does not cite them or go beyond them, for instance by using specialist overlap to estimate *remaining undetected* defects. That would have been the genuinely new move.
7. **Padded.** Paragraphs 4–6 mostly repeat paragraphs 1–3 ("the useful signal is the relationship…", "the scaling question is…").

## Uncertainty
- The tweet itself could not be retrieved. If the posted version differs from the supplied text (for example, it links the CallBook job), points 4 and 5 under "What costs points" would weaken. The grade would rise to 6 at most.
- I could not confirm whether any IMD job class runs a *variable* number of agents. If one does, point 2 under "What costs points" partly falls away.
- I could not check whether SIMD tracks job-level findings data on pages beyond its homepage.

## Unanswered questions
- What is the CallBook job ID, and do its `report.md` and submissions match "24 → 8, one reproduced high-severity issue"?
- Of the 16 findings removed, how many were duplicates, how many were rejected as invalid, and how many were merged as minor?
- Did several of the four *differently-scoped* specialists surface the high-severity issue, or just one? Without this, "independent agreement" is not established.
- Is there any IMD template where agent count varies, so the proposed comparison could actually be run?

## Score rationale
The thesis is competent, specific to IMD, and offers a usable metric list, which puts it above 3–4. It is shallow on mechanism, unverified on its evidence, misdescribes the specialist design, and weak on falsifiability. **Final: 5/10.**
