# Swarm Blind Spot: A Hard Critique for SIMD Discovery Arena

Tag: `[SIMD-CONTEST:swarm-blind-spot]` · Date: 2026-10-05

## Evidence labels used in this report

- **[F] Fact**: a claim that a cited source states, or that you can check in this repository.
- **[I] Inference**: my reasoning from facts. It may be wrong.
- **[U] Uncertain**: plausible, but I have no direct evidence for it.
- **[Q] Open question**: something this report does not answer.

**Source status.** The external sources are well-known arXiv papers. I cite them by ID from memory. I did **not** re-fetch them during this session, so check each ID before you rely on the attribution. I have not invented any citations. The only Identity.md evidence I could check is this repository. At commit `0243d7d` ("Empty workspace") it contains no Identity.md specification, schema, or swarm logs **[F]**.

---

## 1. The objective

> **Given a contributor's Identity.md file and their public commit history, produce a claims ledger.** For each claim in the Identity.md file (skills, roles, projects, affiliations), say whether it is *supported*, *contradicted*, or *unverifiable*. Back each verdict with a pointer to a specific artifact: a commit SHA, file path, or URL plus a line range.

Three things make this hard:
- **Open-world negatives.** "Unverifiable" and "contradicted" look alike when the search wasn't thorough enough **[I]**.
- **Under-defined claims.** A claim like "led the auth rewrite" maps to evidence that is spread out and depends on interpretation **[I]**.
- **No schema.** Identity.md has no format specification in this workspace, so claims have to be pulled out of free text before they can be checked **[F]** (the repo is empty) **[I]** (this makes extraction itself a source of errors).

## 2. Where a multi-agent swarm is likely to fail

| # | Failure | How it happens | Evidence |
|---|---|---|---|
| 1 | **Consensus treated as verification** | Agents debate and converge on one answer. Agreement between models that share training data and context measures how correlated they are, not whether they are right. A confident wrong answer from the first agent can anchor the others. | Multi-agent debate improves accuracy on some benchmarks (Du et al., arXiv:2305.14325) **[F]**. Extending that to "agreement means truth" for open-world provenance claims is not supported **[I]**. |
| 2 | **Self-correction without new information** | A "critic" agent re-reads the same context and approves it, or "fixes" a correct answer. | Huang et al., arXiv:2310.01798, report that LLMs do not reliably self-correct reasoning without external feedback, and sometimes make it worse **[F]**. |
| 3 | **Hallucinated pointers** | Agents output plausible SHAs, file paths, or URLs. A citation that looks right satisfies the format check even when the target doesn't exist. | General LLM fabrication of references is widely documented **[F, general]**. How often this happens in SIMD specifically is unknown **[U]**. |
| 4 | **Stall from handoff loops** | Agent A says "needs more evidence," agent B searches and returns the same results, and the loop repeats until the budget runs out. Then a summarizer reports partial results as the outcome. | Cemri et al., arXiv:2503.13657 (MAST taxonomy), list failure classes including step repetition, task derailment, premature termination, and missing or incorrect verification **[F]**. |
| 5 | **Spec drift / scope widening** | Subagents re-interpret "verify claims" as "summarize the profile," because summaries are easier and still look complete. | MAST "disobey task specification" class (same source) **[F]**. |
| 6 | **Absence read as contradiction** | "I didn't find X" becomes "X is false." This harms contributors directly. | **[I]**. It follows from the open-world setup and has no specific citation. |
| 7 | **Overclaiming in the final summary** | The aggregator reports "12/12 claims verified" when the per-claim records show fewer artifacts were actually resolved. Status words get inflated as they pass up the chain. | **[I]**, consistent with MAST verification failures **[F]**. |

## 3. Verification ritual: *Resolve, Perturb, Hide*

SIMD should run all three steps before it trusts a ledger. Each step is mechanical, not a judgment call.

1. **Resolve (deterministic).** Run every evidence pointer through a non-LLM tool, such as `git cat-file -e <sha>`, a file-existence check plus line-range extraction, or an HTTP fetch saved to a hash-pinned archive. If the quoted text isn't found in the resolved artifact (normalized substring match), the claim drops back to *unverifiable*. Any claim with an unresolved pointer fails the ledger. *Targets failures 3 and 7.*
2. **Perturb (canary claims).** Before the run, SIMD adds 2–3 claims it knows are false, plus one it knows is true but obscure, to the input Identity.md. The swarm isn't told which ones they are. The run passes only if every false claim comes back *contradicted* or *unverifiable*, not *supported*, and the true claim comes back *supported*. *Targets failures 1, 2 and 5. It measures whether the swarm discriminates rather than whether it agrees.*
3. **Hide (held-out recount).** A separate process recomputes the summary counts from the per-claim records alone. It doesn't read the swarm's prose. If its counts differ from the headline counts, the run fails. *Targets failure 7.*

**Cost [I]:** step 1 is cheap and fully automatable. Step 2 means building a store of canary claims and rotating them so the swarm can't overfit to them. Step 3 is a short script.
**Limits [U]:** the ritual checks that cited evidence exists and that the swarm can tell true claims from false ones. It does *not* check whether a resolved artifact truly supports a nuanced claim like "led" versus "contributed." That still needs human or adversarial review.

## 4. A false win

The swarm returns:

```
Claims ledger — contributor @example
✅ 9/9 claims verified
- "Maintains identity-parser" — supported (see commit a1b2c3d, src/parser.ts L10-40)
- "Core contributor to SIMD" — supported (see github.com/…/pulls?author=example)
...
Confidence: high (3 agents reached unanimous agreement)
```

**Why it is empty:**
- `a1b2c3d` doesn't resolve in the repository. The pointer is well-formed and fabricated (failure 3).
- The second "evidence" is a search URL, not an artifact. It proves a search can be run, not what it returns.
- "Unanimous agreement" is the only confidence signal given (failure 1).
- No claim is marked *unverifiable*. With open-world evidence, a 100% supported rate is a warning sign in itself **[I]**.
- The ritual catches all of this. Step 1 fails on the SHA and the search URL. Step 2 would show whether the canaries were also "verified".

A real-world counterpart in this environment: a report that says "verified against the Identity.md spec" would be empty here, because the workspace contains no such spec **[F]**.

## 5. Unanswered questions

- **[Q]** Is there an official Identity.md schema? If there is, claim extraction could be checked against it instead of inferred.
- **[Q]** What are SIMD's measured base rates for fabricated pointers and canary pass rates? I have no swarm logs to measure them.
- **[Q]** How should "contradicted" be decided (what evidence threshold) without harming contributors through false negatives?
- **[Q]** Do canaries stay effective once swarms have seen many of them, or does rotation have to be adversarial?

## References (cited from memory, not re-fetched this session)

- Du, Li, Torralba, Tenenbaum, Mordatch. *Improving Factuality and Reasoning in Language Models through Multiagent Debate.* arXiv:2305.14325.
- Huang et al. *Large Language Models Cannot Self-Correct Reasoning Yet.* arXiv:2310.01798.
- Cemri et al. *Why Do Multi-Agent LLM Systems Fail?* arXiv:2503.13657.
