# Hard grade: "judged acceptance counts failures in the denominator" (@nodeofege)

**Thesis:** <https://x.com/nodeofege/status/2107516363107340564> (posted 2026-10-06 17:00:31 UTC)
**Graded:** 2026-10-06, against live public data fetched about 17:02 UTC the same day.
**Verdict: 6 / 10, flag `thin`. Below the pay bar (8).**

The central observation is true and I could reproduce it. The post stops at the observation: it never finds out what "failed" means, never checks whether the headline number is actually used for anything, and recommends a fix the Explorer already ships.

---

## 1. What the thesis claims

| # | Claim | Type |
|---|---|---|
| C1 | Explorer shows agent #846 at "91% of judged" with 1,764 accepted / 14 rejected / 151 failed | factual |
| C2 | The rate is accepted / (accepted + rejected + failed), so failures sit in the denominator | inference from C1 |
| C3 | Agent #625 shows the same pattern: 1,093 / 9 / 15 → 98% | factual |
| C4 | Rejected and failed are different signals (bad work vs. execution, tooling or environment) | interpretive |
| C5 | If the metric influences reputation, routing or economics, IMD compresses two risks into one number | conditional |
| C6 | Remedy: expose both separately | recommendation |

## 2. Evidence

All sources are public and unauthenticated. Counts are live and drift by the minute.

### Facts (directly observed)

**F1. The tweet text matches the assignment verbatim.** Mirror: `api.fxtwitter.com/nodeofege/status/2107516363107340564`. At fetch time it had 5 views and 0 replies, reposts or likes. The mirror also reports 778 followers, which I did not use in scoring.

**F2. C1 is exactly right.** <https://explorer.imd.fun/agents/846> renders "**1764** accepted · **91%** of judged", and below it Accepted 1764 / Rejected 14 / Failed 151 / Pending 10, "1939 submissions". `GET https://api.imd.fun/seats/846` returns the same: `attempts 1939, accepted 1764, rejected 14, failed 151, pending 10`.

**F3. C3 is partly verifiable.** At fetch time #625 showed 1194 accepted / 9 rejected / 15 failed / 18 pending, displayed "98% of judged". Rejected and failed match the tweet. Accepted had moved from the tweet's 1,093 to 1,194, so the exact figure the author saw cannot be re-observed.

**F4. The Explorer already shows Rejected and Failed as separate counters** on every agent page, and the API returns them as separate fields (`GET /seats/:tokenId`, `GET /seats/records`, documented at <https://imd.fun/docs/>). The thesis's own numbers are read off that separate display.

**F5. Dispatch eligibility is published separately and does not reference the headline rate.** `GET https://api.imd.fun/seats/846/standing` returns two distinct blocks:
- `standing`: `consecutiveFailures`, `pausedUntil`, `breaker: {failures: 3, cooldownMs: 900000}`. This is a failure circuit-breaker: three in a row, 15-minute cooldown.
- `contract`: `good`, `bad`, `probationUntil`, `rules: {bad: 2, of: 3, probationMs: 86400000, capMs: 604800000, windowMs: 2592000000}`. This is a separate good/bad rule with 24-hour probation, a 7-day cap and a 30-day window.

**F6. On-chain reputation feedback is per submission and tagged by verification type, not an aggregate rate.** In the 200 most recent batches from `GET https://api.imd.fun/feedback/batches?limit=200` (2026-10-05 19:46 to 2026-10-06 16:57 UTC), entries were: `review:submission` value 1 ×286, `verification:structural` value 1 ×191, `verification:checks` value 1 ×141, `verification:checks` value 0 ×17. All carry `tag2: acceptance-v2`.

**F7. Every "failed" attempt carries a submission hash.** For #846, 151 of 151 failed work rows have a `submissionHash`; for #625, 15 of 15. No row exposes a failure reason. The job view (`GET /jobs/:id`) shows only the seat that finally got accepted. For example, job `38b6c138-ed1b-43de-83a3-9974b2f1dac1` is "failed" on #846's record and accepted for seat 1536.

**F8. #846's failures are heavily clustered in time.** All 151 are `oracle_assess` attempts. By submission day:

| Day | Attempts | Failed |
|---|---|---|
| 2026-09-26 | 427 | 23 |
| 2026-09-27 | 166 | 0 |
| 2026-09-28 | 36 | 9 |
| 2026-09-29 | 36 | **36** |
| 2026-09-30 | 74 | **74** |
| 2026-10-01 | 8 | **8** |
| 2026-10-02 | 253 | 1 |
| 2026-10-04 | 668 | 0 |
| 2026-10-05 | 271 | 0 |

118 of the 151 failures fall on three consecutive days with a 100% failure rate; the 996 attempts after the last failure have none. Its 14 rejections fall on different days (09-26 ×6, 09-27 ×6, 10-02 ×1, 10-04 ×1).

**F9. Fleet-wide totals** from `GET /seats/records` (714 seats): 913,492 attempts, 878,621 accepted, 8,377 rejected, 10,728 failed, 15,766 pending. The four outcome counts sum to attempts for every seat.

### Inferences (mine, from the facts above)

**I1. C2 is correct, and I tested it harder than the thesis did.** The #846 example alone does not identify the formula: 1764 / 1939 (all attempts, pending included) is 90.97%, which also displays as 91%. I picked eight further seats where the three candidate formulas round to three different integers and fetched their Explorer pages. All eight match accepted / (accepted + rejected + failed):

| Seat | A / R / F / P | a/(a+r+f) | a/(a+r) | a/attempts | Explorer |
|---|---|---|---|---|---|
| 1440 | 898 / 6 / 27 / 19 | 96 | 99 | 95 | 96 |
| 61 | 2129 / 23 / 32 / 34 | 97 | 99 | 96 | 97 |
| 735 | 3560 / 52 / 29 / 56 | 98 | 99 | 96 | 98 |
| 1199 | 2123 / 15 / 148 / 11 | 93 | 99 | 92 | 93 |
| 1498 | 2389 / 11 / 143 / 27 | 94 | 100 | 93 | 94 |
| 138 | 1858 / 22 / 54 / 21 | 96 | 99 | 95 | 96 |
| 70 | 1020 / 10 / 12 / 21 | 98 | 99 | 96 | 98 |
| 174 | 3349 / 23 / 72 / 34 | 97 | 99 | 96 | 97 |

This is still inference. The formula is rendered server-side and I found no published definition of "judged" in the docs or in the client scripts I downloaded.

**I2. The effect is material, which the thesis asserts but never sizes.** Fleet-wide, failures are 56.2% of all non-accepted outcomes (10,728 vs. 8,377 rejections). The fleet rate is 97.87% with failures in the denominator and 99.06% without. Among the 594 seats with at least 100 judged attempts, 34 lose five points or more to failures and 9 lose ten or more. Ranking those seats by the headline versus by accepted / (accepted + rejected) gives a Spearman correlation of 0.556. Treat that figure as rough: most seats sit between 98% and 100%, so ties and small differences dominate it.

**I3. F8 is consistent with C4 for #846.** Three straight days at 100% failure followed by about a thousand clean attempts looks like a systematic cause on that device or its environment, not a drift in answer quality. This is a pattern in timestamps; I have no failure reason to confirm it.

**I4. F7 complicates C4.** A "failed" attempt is not a run that produced nothing: a submission was hashed every time. So "failed" could include submissions that broke a structural or pipeline step after delivery. The thesis's split between "bad work" and "execution, tooling or environment" may not map onto rejected vs. failed as cleanly as it assumes.

**I5. C5's premise looks weak on the public evidence.** The surfaces that would carry reputation and routing (F5, F6) are built on per-submission verdicts, a consecutive-failure breaker and a separate good/bad rule. They already treat failure and quality as different things. The headline percentage appears to be a display figure on a card.

### Uncertain

- Whether the dispatcher or any buyer-side selection also reads the aggregate rate. F5 shows what is published about standing; it does not prove nothing else is used.
- Whether failed attempts are written to the reputation registry at all. My 200-batch sample showed value-0 entries only under `verification:checks`, but oracle work is batched separately (`oracleBatches`) and I did not decode those receipts.
- Which outcome class the `contract.bad` counter counts.

### Unanswered

- What precisely sets an attempt to `failed` rather than `rejected`. No public definition found.
- Why #846 failed every attempt from 09-29 to 10-01.
- Whether "judged" is meant literally. If failed attempts never reach a judge, the label is simply wrong, which is a sharper criticism than the one the thesis makes.

## 3. Claim-by-claim grade

| Claim | Finding |
|---|---|
| C1 | **True.** Exact match to Explorer and API. |
| C2 | **True, but under-argued.** The #846 arithmetic is also consistent with a different formula; the thesis never rules it out. Its #625 example happens to do so, unremarked. |
| C3 | **Plausible.** Rejected and failed match; accepted has since moved, formula still holds. |
| C4 | **Reasonable but unsupported.** Stated as "not necessarily" and "can", with no look at what a failed attempt is. The clustering in F8 would have made the case; the submission hashes in F7 would have tested it. |
| C5 | **Hypothetical, and the available evidence points the other way.** "If this metric ever influences…" is never checked against the public standing and feedback endpoints. The closing "are we pricing…" presumes a pricing role nobody has shown. |
| C6 | **Already done.** The two counters are displayed separately, which is how the author got the numbers. The defensible version is narrower: the headline label and tooltip should say failures are included, or show a second rate. |

## 4. Scoring against the rubric

**For it**
- Concrete, checkable, IMD-specific. Named agents, exact counts, arithmetic shown. This is uncommon among thesis posts and the numbers hold up.
- Honest register: "appears to", "may be", and it concedes the counter-view that buyers care about both.
- The underlying distinction, quality risk versus reliability risk, is the right one to draw, and fleet data shows it is not a rounding matter (I2).

**Against it**
- One observation and two data points. Everything past the arithmetic is conditional or rhetorical.
- No mechanism named. Nothing on what `failed` is, how dispatch pauses a seat, or what gets written to the reputation registry. All of that is one public API call away (F5, F6).
- No quantification of the effect. The fleet-level numbers that would turn a curiosity into an argument are missing.
- The remedy describes the current state of the product.
- The ambiguity in its own lead example goes unnoticed.
- SIMD appears only as a handle. The data is IMD Explorer's; nothing here is about SIMD.
- Ends on a question instead of a finding.

**Why 6 and not 7.** The rubric requires named mechanisms and tradeoffs for 7. There is one tradeoff, stated in a sentence, and no mechanism beyond a division. **Why 6 and not 5.** The claim is falsifiable, specific to IMD, and correct, and it is not a recycled take.

**Originality:** modest. Reading a denominator off a dashboard is a good catch, not a synthesis.
**Depth:** shallow. About 190 words, roughly half of them framing.

## 5. What would have made it pay-grade

1. Rule out the alternative formula with seats where the candidates diverge (I1).
2. Size the effect across the fleet and show the rank reshuffle (I2).
3. Show #846's three-day 100% failure block as the worked case (F8).
4. Check `/seats/:id/standing` and `/feedback/batches` and report that routing and reputation already separate the two risks, then argue precisely about what the headline card misleads on.
5. Propose a specific change: relabel to "of settled", or show accepted / (accepted + rejected) beside a reliability rate.

## 6. Limits of this review

- One snapshot in time; live counts have already moved.
- The rate formula and the meaning of "failed" are inferred from outputs, not from source or a published definition.
- The tweet was read through a third-party mirror because x.com returned HTTP 402 to direct fetch.
- The feedback sample covers about 21 hours and 200 batches.
- Note that `identity.md` the domain is an unrelated site; the project's surfaces are `imd.fun`, `explorer.imd.fun`, `api.imd.fun` and `si-md.xyz`.

```json
{"quality":6,"impactNote":"Flags a real, reproducible labelling issue on IMD Explorer: the 'of judged' headline is accepted/(accepted+rejected+failed), and fleet-wide failures are 56% of non-accepts, so the card blends reliability into what reads as a quality score. Useful prompt for a relabel or a second rate; does not show the number drives routing, reputation or pricing, and public standing/feedback endpoints suggest those already separate the two.","notes":"Strengths: exact, verified figures for #846; correct denominator inference (confirmed on 8 further seats); IMD-specific and falsifiable; concedes the counter-view. Weaknesses: single observation with two data points; never defines 'failed' or checks that failed attempts still carry submission hashes; lead example is also consistent with a different formula; no fleet quantification; 'expose both separately' is already how Explorer and the API present the counts; the routing/pricing concern is purely hypothetical; nothing SIMD-specific; ends on a rhetorical question.","flags":["thin"]}
```
