# Hard grade: Identity.md and SIMD thesis

**Grade: 38/100 — fails as an evidenced account of a controlled experiment.** The useful insight is to distinguish labor supply from demand activation. The fatal leap is to call an evolving subsidy a first-contact-only, controlled trial that isolates price. Neither the treatment restriction nor causal identification is established. Current primary sources also undermine the claim of uniformly full-price payouts.

Assignment: `[SIMD-THESIS]:muxc46bg-rx8fz`. Research snapshot: **2026-10-06 UTC**, with direct API retrievals around 23:54–23:55 UTC. The submitted thesis is undated and ends mid-word at “durab”; this report grades the supplied text, without reconstructing its ending. Current discrepancies do not prove what held at an unspecified earlier date.

This is a grade of the argument and its evidence, not a claim that either project has failed commercially. Missing evidence earns no credit, but is not evidence of zero demand.

| Dimension | Score | Reason |
|---|---:|---|
| Factual grounding and attribution | 12/30 | Price and substantial accepted work are supported; vault claims are unreconciled and the thesis supplies no dated sources. |
| Causal reasoning | 3/25 | No documented control, assignment mechanism, stable treatment, or first-use enforcement. |
| Measurement and falsifiability | 12/20 | Numerical targets are useful, but denominators, maturity, funding attribution and decision timing are underspecified. |
| Economic reasoning | 6/15 | Correctly questions subsidized demand; conflates repeat payment with viability and ignores full production costs. |
| Calibration and clarity | 5/10 | Separates two questions well, then repeatedly presents hypotheses as established mechanism. |
| **Total** | **38/100** | **Major revision required.** |

The rubric is this reviewer's explicit judgment, not a calibrated statistical score.

## What the evidence supports

Labels used below: **observed** means retrieved from a named public source; **operator claim** means a project's assertion, not independent verification; **inference** is this reviewer's reasoning; **unverified** means the evidence obtained cannot establish the claim.

| Thesis claim | Finding and attributable evidence |
|---|---|
| NFT-gated agents; 2,000 capped seats | **Partially supported.** Identity.md documents device-to-seat pairing and holder authorization. The 2,000-seat figure is reported by Bankless; this review did not independently inspect the collection's mint cap or mutability. A seat cap is not a cap on compute, throughput, or independent operators. [IMD documentation](https://imd.fun/docs/), [Bankless, September 25 snapshot](https://www.bankless.com/read/inside-imd-ethereum-s-new-ai-swarm-experiment) |
| Anyone can open a job for 0.5 IMD | **Price observed.** The live `job.open` capability reports 500000000000000000 atomic units, 18 decimals, Ethereum mainnet, and IMD contract `0xd34a99bc0f67ae1bbd63c660e6d0b0dd03e263b7`: exactly 0.5 IMD. This is a service price, not a guarantee of admission, successful completion, or independently correct output. [Live capabilities](https://api.imd.fun/requests/capabilities), [saved response](evidence/capabilities.txt) |
| Every job, submission and failure is publicly recorded | **Visible records supported; “every” unverified.** The explorer exposes jobs and failures. Documentation exposes seat outcome totals. Neither proves exhaustive retention or that all work events are blockchain records. A public HTTP explorer is not itself an immutable ledger. [Jobs explorer](https://explorer.imd.fun/), [IMD documentation](https://imd.fun/docs/) |
| Individual agents exceed 90% on large samples | **Supported as platform-recorded acceptance.** Seat 355 has 4,034 acceptances in 4,291 attempts (94.01%); seat 527 has 4,120 in 4,289 (96.06%). This is not evidence of thousands of distinct satisfied paying customers. [Seat records](https://api.imd.fun/seats/records), [saved response](evidence/seats.txt) |
| SIMD observes rather than operates agents | **Supported narrowly.** Its published frontend docs describe an observer/tooling surface that reads IMD and does not pair devices. Hire does pay to open jobs. This does not establish that people controlling SIMD operate no seats elsewhere. [SIMD published frontend](https://www.si-md.xyz/app.js?v=66), [saved docs excerpt](evidence/simd-docs-excerpt.txt) |
| Trading fees enter the named vault and fully reimburse jobs | **Operator-described funding mechanism; full capital path unverified.** The API identifies the thesis's vault and token relationship, but neither incoming fee provenance nor one-to-one job settlement was independently reconciled. Its actual reported payout rows conflict with a universal full-price characterization. [SIMD state](https://www.si-md.xyz/api/state), [saved vault extract](evidence/simd-vault.json) |
| Repeated 500–1,300+ IMD balances and consecutive 100% batches | **Historical series unverified.** No dated balance/eligible-job series was obtained. A present balance does not validate historical ranges or batch coverage. |
| Only initial trial is funded | **Not established; contrary indications.** SIMD's docs describe payer-funded Hire and later refunds, with no first-wallet-only treatment rule in the inspected sections. A social mirror attributes broader reimbursement language to the operator. [SIMD frontend](https://www.si-md.xyz/app.js?v=66), [operator posts via Sotwe](https://www.sotwe.com/SuperIMD_eth) |
| Acceptance remained high under reimbursement | **Unverified longitudinal claim.** Current cumulative counts provide neither a pre/post comparison nor a reimbursement-linked cohort. |
| Independent wallets completed and shipped useful work | **Unverified as stated.** Distinct addresses do not establish distinct humans or economic independence. Transfers do not prove shipped artifacts, customer use, or reimbursement eligibility. |

The docs-derived claims above are limited to the documented interface; no worker registration, payment, job creation, or wallet connection was performed.

## The strongest factual problem: policy labels do not match payout evidence

The first-party [SIMD state API](https://www.si-md.xyz/api/state), retrieved at 23:55:05 UTC, contains a vault reading timestamped **2026-10-06T23:53:56.679Z**. Its [saved extract](evidence/simd-vault.json) reports:

| Field | Observed API value |
|---|---|
| Vault | `0xd60483Eb8004e3DE3e283b3efF0e67FBb57f9B21` |
| Balance | 1,658.034729 IMD |
| Policy display | `share: "100%"`, `payoutAmount: "0.5 IMD"` |
| Aggregate labels | `distributed: "36.75 IMD"`, `jobsPaid: 145`, `fullPrice: 0`, `recipients: 43` |
| Returned payout rows | 40; each 0.25 IMD; each `matchesShare: false` |

The newest row identifies transaction `0xd9373a85bd825e2488d6464a0efa56f430572b414a37aca9e731675cd5193c56`, at 23:42:11 UTC, paying 0.25 IMD to `0x5b92D2b38AB372bFf63f55113b98253aDFF62A74`. Its [Etherscan receipt link](https://etherscan.io/tx/0xd9373a85bd825e2488d6464a0efa56f430572b414a37aca9e731675cd5193c56) was inaccessible to this review; it is a verification target, not an independently inspected receipt.

**Interpretation:** the project's own data contradicts the blanket presentation of payout rows as clean 0.5 IMD transfers. Forty rows are not necessarily forty jobs or forty transactions: some transaction hashes repeat for different recipients. The aggregate labels have no established lifetime scope here. Do not convert them into a historical coverage percentage. Split payments, different eligibility rules, stale indexing or accounting bugs remain possible explanations; none was established.

The [published SIMD frontend documentation](https://www.si-md.xyz/app.js?v=66) separately describes a hot payer wallet, reimbursement over time and possible payer depletion before refunds arrive. It labels its full verified-task-price policy **“NOT IN FORCE”** pending an on-chain release rule. This is documentation evidence, not a deployed-contract audit. It does not prove that discretionary full payments never occur. It does make “precisely as designed,” guaranteed full settlement and absence of delay unjustified. [Preserved excerpt, source lines 2990–3016](evidence/simd-docs-excerpt.txt)

The [Hire summary endpoint](https://www.si-md.xyz/api/hire/sent?summary=1) identifies payer `0x9fAdAB91f6FA03DBd7F4F8a08A338704bAAcf63f` and returns count 100. That is a payer-order summary, not 100 customers; IMD documents its source route as capped at 100 orders. A pooled payer can hide end-user identity. [Saved response](evidence/simd_sent.json), [IMD route documentation](https://imd.fun/docs/)

For context only, the [Sotwe mirror](https://www.sotwe.com/SuperIMD_eth) attributes all-job reimbursement statements, walletless hiring, 121.5 IMD over 386 refunded tasks, and a 1,341 IMD balance to the operator. Earlier search material instead reported 14.75 IMD over 30 jobs. These are changing, indirectly retrieved promotional snapshots with relative timestamps. They do not reconcile with the API by themselves. The mirror also mentions a one-job-per-linked-X-account limit without enough context to equate it to a lifetime first-wallet restriction. None is used as technical proof.

## High acceptance is real telemetry, not product-market fit

The saved [seat API snapshot](evidence/seats.txt) contains 728 records totaling 913,813 attempts: 878,892 accepted, 8,379 rejected, 10,762 failed and 15,780 pending. These categories sum correctly. Calculated acceptance is **96.18% of all attempts**, or **99.06% of accepted-plus-rejected outcomes**. Changing the denominator materially changes the headline.

This supports functioning work infrastructure and high recorded acceptance. It does not establish independent grading accuracy, useful output, unique paid jobs, or external customer demand. Multiple agents, retries and work stages can contribute observations without creating new customers. Selection of the best agents also cannot establish typical performance. Platform judgment must be supplemented by blinded output review and customer outcomes.

The [jobs explorer](https://explorer.imd.fun/) visibly includes thesis-grading tasks alongside development work. That is evidence of heterogeneous workload, not a measured share of organic demand. This report is itself generated through the contributor assignment being examined; its completion is not independent proof of the thesis's economic claims.

## Why this is not an isolated experiment

These are **methodological inferences**, not additional claims about hidden implementation:

1. **A subsidy is a treatment, not a control.** Without randomized eligibility or a credible comparison design, conversion cannot be attributed to reduced price. Users choosing a promotion already differ from nonparticipants.
2. **The treatment bundles several changes.** Walletless hiring changes signing friction and onboarding as well as monetary cost. Promotion, novelty, token incentives, changing quality and shifting reimbursement rules can also move demand.
3. **Reimbursement is not zero upfront cost.** Paying and later reclaiming money retains liquidity requirements, waiting time and reimbursement uncertainty. A sponsor-paid job is a different treatment.
4. **Wallets are imperfect units.** One person can rotate wallets; several users can share the SIMD payer. “Never reimbursed” wallets may still spend prize money, transfers from affiliates or prior subsidized earnings.
5. **Repeat paid use is descriptive evidence, not incremental lift.** Some converters would have paid anyway. Nonconversion can reflect infrequent need or bad timing, not necessarily poor quality.
6. **Trading-fee finance is not customer finance.** Speculative trading may support subsidies independently of customer value. Reimbursement sustainability and unsubsidized demand are separate outcomes.

The statement that price is the “exact variable” determining a self-sustaining labor market is therefore false as a general economic proposition. Reliability, task fit, frequency of need, acquisition, capacity, operator costs and price all matter. Nor does IMD's potential value logically depend exclusively on this particular SIMD experiment succeeding.

## The four thresholds are proposals, not results

No retrieved cohort dataset establishes any threshold. **All four are presently unmeasured in this review**, not zero and not passed. There is no identified trial start date or preregistered operator adoption of these targets.

| Proposed target | Problem | Minimum defensible definition |
|---|---|---|
| 30% convert within 45 days | Immature users bias the denominator; first-time wallet is not necessarily first-time customer. | Freeze enrollment cohorts. Count economically distinct eligible first users with complete 45-day follow-up; numerator is at least one settled, admitted, genuinely unsubsidized subsequent job. Publish counts and uncertainty intervals. |
| Median conversion time ≤21 days | A fast handful can pass while almost everyone never returns. Undefined if nobody converts. | Report conditional median among converters alongside overall conversion and censored time-to-event results. Define time zero and whether “second self-paid job” means the next job overall or the second paid follow-up. |
| Fees cover ≥75% across rolling 10 days without sustained drawdown | Ratio of sums versus average daily ratios is unspecified; zero-spend days and extra deposits can distort results. At exactly 75%, an otherwise closed vault loses 25% of reimbursement spend. | Use verified fee inflow divided by reimbursement outflow over every complete 10-day UTC window; publish both sums. Separate donations, prizes, operational spending, payer float and unpaid obligations. Define a numeric drawdown tolerance and duration. |
| 40 never-reimbursed paying wallets within 90 days | Can count existing users or Sybil wallets and does not identify subsidy-induced acquisition. | Define new-user baseline and observation horizon; exclude controlled/sponsored payers and known subsidies. Treat this as breadth of observed paid usage, not causal conversion. |

“Any two failures for a full calendar month” mixes 10-, 45- and 90-day clocks and lacks a cohort evaluation rule. A cohort enrolled at the end of a month cannot have a full 45-day outcome that month. Predefine decision dates after relevant follow-up matures, minimum sample sizes, treatment of insufficient data, and whether failure means point estimates or uncertainty bounds. Overlapping rolling windows are correlated, not independent replications.

To test first-contact cost credibly: prospectively assign eligible new users to a one-time subsidy or the normal price, keep onboarding and service access comparable, enforce and record the one-time rule, and follow both groups for the same duration. Estimate incremental paid retention, not just treated conversion. If randomization is unavailable, label the study observational and specify its identification assumptions. Quality changes require time/version tracking and comparable task mix.

Even passing all four thresholds would not establish profitability. Measure customer revenue net of subsidies against inference, hosting, review, gas, support, failure/rework and acquisition costs; distinguish token rewards from sustainable cash earnings.

## Unanswered questions and evidence needed

- What dated, authoritative policy defines eligibility, first-use limits, refund amount and termination? Are direct refunds and walletless Hire separate programs?
- Which fee collector and contracts fund the vault, at what share, and through what conversions? Who can redirect or stop funds? No fee-provenance trace or control audit was completed.
- Can every claimed reimbursement be joined to an order, actual payer, job ID, amount, receipt/log index, eligibility decision and shipped artifact? Include unpaid eligible jobs to avoid survivor bias.
- Why do current 100% labels coexist with 0.25 IMD payout rows and different promotional totals? Publish block-pinned reconciliation rather than choosing the most favorable dashboard number.
- Which users return after subsidies stop, and who bears their cost? Publish mature cohorts, exclusions, related-wallet assumptions and anonymized end-user mapping where a shared payer is used.
- What fraction of accepted outputs pass independent review and get used? What is the net cost per useful completed job?

The defensible replacement thesis is: **Identity.md exposes a paid agent-work service with substantial recorded acceptance. SIMD adds observation, sponsored access and reported vault payments. Whether those subsidies create incremental, durable, unsubsidized demand remains open. Current evidence does not establish a controlled first-contact-only trial, uniformly full-price settlement, or a self-sustaining market.** Near-zero conversion may be a conservative scenario; it is not an empirically demonstrated base-rate forecast here.

## Research limits and reproducibility

Primary sources were the IMD docs, explorer and direct API responses, plus SIMD's served frontend and state endpoints. Secondary sources were used only with explicit qualification. Search results for unrelated Solana SIMDs and the differently named IMD staking share were excluded. SIMD's own site identifies token `0xbb0c1f82a2ea0253ea3d91c2f05caded82133415`; that identity is distinct from IMD. [SIMD site](https://www.si-md.xyz/), [saved HTML](evidence/simd_site.txt)

Etherscan address and transaction pages were inaccessible through the browser tool; the Blockscout token-transfer API returned HTTP 403. Direct Python HTTP requests successfully retrieved the project APIs after browser access failed. **No independent blockchain receipt, full transfer history, historical balance series, user cohort, or deployed contract was verified.** API timestamps are not pinned blocks. Failure to retrieve evidence is not evidence of absence.

[Retrieval metadata](evidence/retrieval.json) records direct URLs, times and errors. The large SIMD state payload was reduced to its exact vault subtree and top-level timestamp; unrelated fields were discarded. The frontend excerpt includes source line numbers and the full downloaded script's SHA-256. [Offline calculation/check script](check_report.py) verifies saved arithmetic and evidence consistency, not source truth. No external reviewer certified this report.
