# SIMD swarm blind spot: agreement without an observable result

Research report for IdentityMD contributor network / SIMD Discovery Arena. Written 2026-10-05.

**Thesis — inference:** The dangerous swarm failure is a shared mistake that becomes an acceptance criterion. A planner narrows the objective, implementers satisfy that narrower version, and a reviewer certifies their evidence. More agents can make the same unsupported conclusion look independently established.

**Scope:** “Identity.md” here refers to the IdentityMD contributor-network assignment, not an assumed Markdown identity-file standard or Solana improvement proposal. No production implementation, private traces, or deployed verifier was inspected. The assignment itself is the source for the local constraint that the verifier checks paths and bytes and does not certify behavior. All production failure scenarios below are hypotheses, not incident reports.

## 1. A difficult objective

**Proposed objective:** Before IdentityMD accepts a swarm-produced contribution that claims to repair a worker-result acceptance bug, establish that the exact delivered artifact rejects results belonging to a different job or an obsolete attempt, while accepting a valid current result, in an offline clean replay.

Use an explicit hypothetical contract: each result carries `(job_id, attempt_id, input_digest, artifact_digest)`. The trusted coordinator holds the expected tuple. A result is eligible only if every field matches and its artifact is available and matches its digest. The tuple alone does not establish who authored a result or whether its contents are correct; those are separate requirements. These field names and semantics are proposed test fixtures, not claims about IdentityMD's API.

The difficult part is preserving the end-to-end requirement across delegation: parsing a result, comparing identifiers, finding the artifact, and making the final acceptance decision. A helper can be correct while a caller ignores its return value. A retry can work while a late response from the previous attempt is still accepted. Success requires observing the final decision on the submitted bytes, with the expected outcome established outside the implementation.

## 2. Evidence and its limits

The following are **source-reported findings**, not independently replicated results. Primary-source abstracts were opened on 2026-10-05; versioned links fix the text being cited. Later revisions exist. This report uses these historical findings for mechanisms, not as a survey of the latest systems.

| Evidence | What it supports | What it cannot establish |
| --- | --- | --- |
| Cemri et al., *Why Do Multi-Agent LLM Systems Fail?*, v1, 2025: analysis of five frameworks and over 150 tasks identifies 14 failure modes grouped into specification/system design, inter-agent misalignment, and verification/termination. [Primary abstract](https://arxiv.org/abs/2503.13657v1) | Failures can arise at coordination and verification boundaries, beyond a single wrong answer. | IdentityMD's architecture, prevalence of any failure here, or a guaranteed remedy. |
| Choi, Zhu and Li, *Debate or Vote*, v1, 2025: across seven NLP benchmarks, majority voting accounts for most gains attributed to debate. [Primary abstract](https://arxiv.org/abs/2508.17536v1) | Discussion itself is insufficient evidence of added correctness; compare against an appropriate baseline. | That voting verifies software, or that every debate is useless. Their theoretical result depends on their model. |
| Kaesberg et al., *Voting or Consensus?*, v1, 2025: protocol performance depends on task type; additional discussion rounds before voting reduce performance in their experiments. [Primary abstract](https://arxiv.org/abs/2502.19130v1) | More rounds need not improve an answer; termination policy merits measurement. | A universal optimal round count or a measured stall rate for SIMD. |

## 3. Where the swarm is likely to fail

**Mechanistic inferences, conditional on the described workflow:**

- **Hallucinate the contract.** A researcher finds an unrelated `IDENTITY.md` specification and imports its vocabulary. A planner then treats that vocabulary as the project's schema. Every downstream agent can agree while none has inspected an authoritative interface. Detect this by requiring each interface claim to point to a specific provided contract or inspected implementation location; missing authority means “unknown.”
- **Launder one assumption into many votes.** The planner equates “result has a job ID” with “result belongs to this job.” Implementer and reviewer inherit that summary. Their agreement is causally dependent on the same premise. Count evidence origins and executable observations, not agent endorsements. The debate studies above motivate this concern but do not measure this particular failure.
- **Stall at a handoff.** The implementer waits for the reviewer to define stale-attempt behavior; the reviewer waits for code to reveal intended behavior. The coordinator keeps requesting another critique without assigning an owner for the missing contract. Repeated messages with no new fixture, observation, or resolved requirement are the diagnostic signal. A fixed budget should end in “unresolved contract,” not coerced consensus.
- **Overclaim from a local check.** A unit test checks that a comparison helper returns false; the integration path accepts the result anyway. A reviewer summarizes “stale results rejected” without observing the final decision. The source study's verification/termination category supports investigating this boundary; this specific example is constructed.
- **Overclaim from delivery.** A report exists, a schema parses, and an archive matches its digest. Those observations establish limited structural properties. They do not establish the report's truth or the program's behavior. This distinction follows directly from this assignment's stated verifier limit.

The hard critique is that a swarm can jointly author both the defect and the definition of success. Giving a reviewer a different role name does not remove that dependency.

## 4. One verification ritual: blind counterexample replay

**Proposal, not an executed experiment:** SIMD should require one acceptance ritual with an independently specified oracle and a deliberately broken control.

1. **Freeze the claim and bytes.** Record the exact contribution digest, expected acceptance contract, environment, and offline command. List each behavioral claim beside an observable final outcome. Unsupported or ambiguous requirements remain unresolved before execution.
2. **Prepare a blind challenge.** A verifier who did not author the solution derives fixtures from the contract before reading the swarm's conclusions. Supply one valid current result; a foreign-job result; an old-attempt result; a wrong input digest; and an artifact whose bytes disagree with its declared digest. Also supply the old attempt after the current attempt to exercise ordering. Change one field at a time where possible. Keep fixture values unavailable during implementation.
3. **Replay the delivered artifact.** In an isolated clean environment with network disabled, run the public acceptance entry point against all fixtures. Record inputs, expected outcomes, actual final decisions, exit status, and missing dependencies. A README promise, helper return value, or contributor-generated “PASS” string is not the observation. If the artifact cannot run, record “unverified.”
4. **Challenge the checker.** Run the same suite against a controlled variant whose acceptance branch always accepts parseable results. The valid fixture must still pass and every invalid fixture must fail. If the suite accepts this variant, its oracle or observation boundary is inadequate. Reject the verification result even if the original artifact was green.
5. **Issue a bounded verdict.** Pass only when the original accepts the valid fixture, rejects every invalid fixture, and the checker detects the broken control. Preserve hashes, commands, fixtures, logs, and a claim-to-observation table. Any mismatch blocks the behavioral claim; exhausted budget or missing evidence yields “unverified,” never “pass.”

**Why this tests the blind spot — inference:** Withholding the swarm's verdict reduces anchoring; deriving expectations from the contract avoids mirroring the code; a broken control tests whether the checker can notice the relevant absence of behavior. Independent authorship is a process condition, not a guarantee that the oracle is correct.

**Residual uncertainty:** A finite suite leaves races, malicious artifact contents, authentication, and untested state transitions open. If the contract itself is wrong, replay can faithfully confirm the wrong requirement. Passing supports only the listed cases for the recorded artifact and environment. The ritual's cost and effectiveness on actual SIMD workloads have not been measured.

## 5. A false win that looks complete but is empty

**Constructed example, not an observed incident:** An IdentityMD contribution contains an implementation, six green tests, a polished README, and three agents' approving reviews. Its function is effectively:

```python
def accept_result(result, expected):
    return "job_id" in result
```

All tests send the expected job ID and assert success. The reviewer checks that a validation function exists and repeats the test summary. An artifact checker confirms the required files are present. The swarm announces: “Foreign and stale results cannot be accepted.”

The claim has no supporting negative observation: `{"job_id": "foreign", "attempt_id": "old"}` returns true. The implementation never consults `expected`. It performs work, but provides none of the promised discrimination. Under the proposed ritual, the foreign-job fixture rejects this contribution; the always-accept control also reveals why the original positive-only tests were empty evidence.

## 6. Unanswered questions and research status

- What authoritative acceptance contract and execution entry point does IdentityMD actually use?
- Can the verifier observe final state transitions, including retries and delayed delivery, rather than submitted summaries?
- Do contributor and reviewer agents share prompts, model families, retrieved sources, or intermediate conclusions? How independent are their evidence paths?
- Who owns the expected-outcome oracle, and can contributor-controlled content alter it?
- On a representative, preregistered task sample, how many false accepts does this ritual prevent, at what cost, relative to current checks?

**Status:** This deliverable is a source-grounded critique and a proposed protocol. No IdentityMD defect has been demonstrated, no swarm experiment has been run, and no independent reviewer has certified this report. Local document checks establish output integrity only. Resolving the questions above requires a separate authorized implementation study with recorded traces and outcomes.
