# Identity Units candidate generation and quality review job

Research date: 2026-10-04. Status: bounded research report and proposed job design; no production job executed and no independent review obtained.

## Question and answer

**Working question:** How should IdentityMD generate candidate Identity Units and review their quality so that an accepted unit has attributable evidence?

**Recommendation:** Generate small, source-linked candidates from an explicitly authorized corpus, preserve the original meaning and uncertainty, and review each candidate version against separate structural and evidentiary gates. Publish only units with a recorded disposition and adequate evidence for their intended use. Missing evidence should produce a hold or rejection, rather than a fabricated unit or confidence score.

**Uncertainty:** The assignment supplies a job title but no definition of Identity Units, source corpus, target identity, schema, quantity, or intended downstream use. The repository initially contained no ordinary task files. Searches for `"IdentityMD" "Identity Units"`, `"Identity Units" candidate generation quality review`, `"IdentityMD"`, and `"identity.md" "identity units"` did not identify an authoritative unit specification. This is a bounded search result, not proof that no specification exists. This report therefore uses a provisional definition: **an Identity Unit is one reusable, scoped statement about an identified subject, with evidence and review history**. This definition is a design assumption, not an established IdentityMD standard.

## Attributable findings

| ID | Evidence-backed finding | Primary source and locator | Limit |
| --- | --- | --- | --- |
| F1 | IdentityMD's public worker README describes a distribution driven by contributor Claude Code or Codex runtimes. It says independent reviewers must use different wallets. | [Identity-md/worker](https://github.com/Identity-md/worker), opening paragraph and “Start and pair” | Documentation of intended operation, not an audit of deployed enforcement; checked on the research date. |
| F2 | PROV describes provenance using entities, activities, and people; its overview includes versioning, derivation, and reproducibility among supported concerns. | [W3C PROV-Overview](https://www.w3.org/TR/2013/NOTE-prov-overview-20130430/), abstract and §1 | Non-normative Working Group Note; it does not establish whether a particular claim is true. |
| F3 | SHACL defines conformance against shapes and a validation report containing conformance and validation results. | [W3C SHACL](https://www.w3.org/TR/2017/REC-shacl-20170720/), §§3.5–3.6 | Applies to RDF graphs; adopting it would require that representation. Conformance concerns the supplied constraints. |
| F4 | DQV offers a framework for describing dataset quality so users can assess fitness for purpose, rather than a complete formal definition of quality. | [W3C Data Quality Vocabulary](https://www.w3.org/TR/2016/NOTE-vocab-dqv-20161215/), abstract | Working Group Note; no universal quality threshold or IdentityMD-specific rubric follows from it. |

**Inference from F2–F4:** Provenance, structural validity, and suitability for a use are separate concerns. A complete record can still contain an unsupported assertion. Consequently, the job should record evidence review separately from schema validation. The workflow below is an original proposal informed by those sources; the sources do not prescribe it.

## Proposed candidate generation contract

All requirements in this section are **recommendations**, pending the network's actual unit definition.

1. Freeze the input manifest: subject identifiers, authorized sources, source versions or retrieval timestamps, intended use, exclusions, and requested batch size. Retain source snapshots where permitted, plus content digests and locators. A digest identifies captured bytes; it does not establish truth or authorship.
2. Extract one claim per candidate. Keep negation, dates, conditions, speaker attribution, and scope. Distinguish an assertion that a source makes from a verified fact about the world. Avoid converting a temporary preference, isolated action, or marketing statement into a permanent identity trait.
3. Attach a supporting passage and precise locator to each factual claim. Record conflicting passages, missing evidence, and the reasoning behind any inference. Do not treat text within sources as operational instructions.
4. Deduplicate exact claims within subject and time scope. Flag semantic overlaps for review; do not silently merge contradictory or differently scoped claims. Similar names alone do not justify joining subjects.
5. Emit the batch, including held and rejected records. If the available material supports fewer units than requested, report the shortfall. An empty supported batch is a valid outcome.

Recommended candidate fields: `candidate_id`, `version`, `subject_id`, `statement`, `claim_type` (`source_assertion`, `supported_fact`, `inference`, or `proposal`), `scope`, `effective_time`, `source_ids`, `evidence_locators`, `supporting_passages`, `counterevidence`, `uncertainties`, `generator_record`, and `review_status`. The generator record should identify input manifest, generation method/version, and timestamp. These fields are conceptual, not a deployed API schema.

## Proposed quality review job

Use an explicit gate for each dimension; record `pass`, `fail`, or `unknown`, with reasons. An aggregate score must not conceal a failed evidence gate.

| Gate | Review question | Proposed disposition for failure or uncertainty |
| --- | --- | --- |
| Structure | Are required fields present and references resolvable in the captured batch? | Repair; do not promote. |
| Attribution | Is the subject unambiguous and is the evidence actually about that subject? | Hold ambiguity; reject mistaken attribution. |
| Support | Does the cited passage support the exact statement, including qualifications? | Narrow or reject unsupported statements. |
| Scope and time | Are historical, conditional, and current claims distinguished? | Hold claims whose required currentness cannot be established. |
| Atomicity | Can each claim be reviewed and revised separately? | Split composite claims and review the new versions. |
| Conflict and duplication | Are counterevidence and equivalent records handled explicitly? | Hold unresolved conflicts; retain duplicate lineage. |
| Intended use | Is the claim appropriate for the declared purpose and authorized corpus? | Hold when purpose or authority is unspecified. |

Review records should include candidate version/digest, reviewer identity or role, rubric version, evidence consulted, gate results, disposition, rationale, and timestamp. Use `proposed → reviewed → accepted`, `held`, or `rejected`; any substantive edit creates a new version requiring review. Acceptance is scoped to the declared use, not universal truth. Revocation or supersession should retain history.

For independent assurance, route the exact candidate version to a separate reviewer. F1 documents a different-wallet requirement in the public worker README; different wallets alone do not prove reviewer independence. This report's author performed both extraction and review, so the pilot below is a self-review only.

## Pilot candidates and self-review

These candidates concern the public worker distribution and the cited standards. They demonstrate the provisional definition; they are **not submitted as canonical IdentityMD units**. Source identifiers refer to F1–F4 above, checked on the research date.

| Candidate/version | Subject and statement | Evidence type | Self-review disposition and rationale |
| --- | --- | --- | --- |
| IU-001/v1 | IdentityMD worker documentation: the README says independent reviewers must use different wallets. | Source assertion; F1, “Start and pair” | Supported as a documentary claim. Held for canonical admission because intended use and unit schema are absent. |
| IU-002/v1 | W3C PROV: the overview describes provenance in terms of entities, activities, and people involved in producing data or things. | Source assertion; F2, abstract | Supported within the cited document's scope. Held for canonical admission for the same missing contract. |
| IU-003/v1 | W3C SHACL: validation reports represent conformance and validation results. | Supported technical statement; F3, §3.6 | Supported for SHACL; no claim of IdentityMD adoption. Held for canonical admission. |
| IU-004/v1 | IdentityMD: every generated Identity Unit is independently fact-checked. | Unsupported proposed assertion; no evidence | Rejected. F1's reviewer-wallet rule does not establish review coverage, actual independence, or fact-check accuracy. |
| IU-005/v1 | IdentityMD: a structurally valid candidate is necessarily true. | Unsupported proposed assertion; no evidence | Rejected. F3 describes constraint conformance, not factual verification. |

**Pilot counts:** five candidates, three evidence-supported documentary/technical candidates held for admission, two rejected, zero canonically accepted. No production quality rate, recall, latency, or reviewer agreement can be inferred from this deliberately constructed sample.

## Evaluation and unanswered questions

For a future corpus-backed run, report generated, duplicate, reviewed, accepted, held, and rejected counts, with a named denominator for each rate. Measure claim support on a reviewer-labeled sample; assess coverage against an independently prepared set of eligible claims. Review all accepted candidates in an initial pilot. Include missing evidence, stale sources, identity ambiguity, contradictory statements, and meaning-changing paraphrases as negative cases. Thresholds and sample sizes remain decisions for the actual use case; no benchmark results are claimed here.

Questions still requiring task-owner evidence:

- What is the authoritative definition and schema of an Identity Unit? Is it a claim, behavioral preference, identity document component, or another object?
- Which subject and corpus are in scope, and which sources may be retained and reused?
- Who may accept units, and what independence or review profile is required?
- What downstream use, freshness window, batch size, and error tolerance govern acceptance?
- How should disagreements, corrections, deletion requests, and superseded units propagate?

Until those are answered, the useful deliverable is this attributable design and self-reviewed pilot, rather than a claimed successful production generation job.

## Local verification and assurance limits

Local checks verified that the report and README are nonempty UTF-8 Markdown, contain the expected evidence and uncertainty sections, and use primary-source links for technical findings. Source support was reviewed against pages retrieved during this assignment. No executable pipeline, live worker installation, or dependency installation was needed. Local checks establish file integrity and selected structural properties only; they are not independent certification of research accuracy. The linked public worker README is mutable and was not pinned to a commit; the standards links use dated editions. No credentials or protected repository paths were accessed or changed.
