**Hackathons judged by AI agents: evidence-backed inventory and how judging worked** Research cutoff and access date: **30 September 2026**. Scope: worldwide, publicly discoverable evidence, primarily English-language sources. **Answer.** AI agents have already evaluated real hackathon entries, but “AI-judged” covers substantially different arrangements. The strongest documented examples found are Cognizant’s Vibe Coding event, NEBULA:FOG:SINGULARITY, Synthesis, four HackerRank Orchestrate editions, Cursor Boston × PyData, Code Forward, the Devin × Claw Collective × Qwen hackathon, and one judge’s workflow at SANS FIND EVIL!. Most of these retained human involvement. The evidence and qualifications for each appear below. This is **a documented inventory, not a provably exhaustive list of every such hackathon**. Private events, unindexed pages, deleted announcements, and undisclosed use by individual judges cannot be enumerated reliably. No global registry was identified. Searches included earlier years; failure to find a qualifying early event does not establish when AI judging first occurred. **What counts.** An AI must evaluate entries, generate consequential scores or shortlists, or assist an actual judge. Building AI agents as the hackathon challenge does not qualify by itself. Ordinary unit-test grading and public upvotes do not establish AI-agent judging. “Agent” is used broadly where sources describe an evaluator that reviews artifacts, interviews entrants, uses tools, or participates in an orchestrated panel. Where only an unspecified AI scorer is documented, that limitation is explicit. **Evidence labels.** “Documented” means an organizer, system builder, or participating judge reports actual use, preferably with results; it does not mean independently audited. “Specified” means published rules describe the mechanism but completed judging was not established. “Uncertain” means material identifying or implementation details remain missing. All mechanisms below are source-reported facts unless labeled **inference**; unknowns are not filled in from analogous systems. **Inventory of documented deployments** | Event or series | When | Actual role of AI | Evidence boundary | | --- | --- | --- | --- | | Cognizant Vibe Coding | August 2025 | Multi-agent screening; humans chose finalists/winners | Organizer’s retrospective; details below | | NEBULA:FOG:SINGULARITY | 14 March 2026 | Live multimodal judging, ensemble scoring and comparative ranking | Public AI results; final prize authority not fully explained | | Synthesis | March 2026 | Partner-specific agent reviewers proposed decisions to humans | Organizer’s retrospective explicitly preserves human decision authority | | HackerRank Orchestrate | May, June, August, September 2026 | Artifact evaluation plus AI voice interview | Four separately documented editions; model/version and override details incomplete | | Cursor Boston × PyData Data Science Hack | 13 May 2026 | Blind AI agent selected notebook awards | Distinct from the human presentation awards | | Code Forward, Money Forward India | 2026, before 2 July retrospective | Five-persona panel supplied scoring, coaching and judging drafts | Human judge retained judgment | | Devin × Claw Collective × Qwen, #BuildForMsia | 23 August 2026, Kuala Lumpur | AI jurors reviewed submissions and questioned teams on Discord | Builder reports deployment on 40+ cases; final-round authority unclear | | SANS FIND EVIL! | 2026; judge’s account dated 28 August | One human judge used an adversarial multi-agent review workflow | Evidence of that judge’s practice, not an event-wide autonomous jury | The table represents **eight event families and eleven editions** when the four Orchestrate editions are counted separately. It is a count of the core evidence set, not a worldwide total. Additional AI-scoring cases, announcements and experimental rounds are retained below rather than silently discarded. **Cognizant Vibe Coding — specialized agents screen at corporate scale.** The company’s August 21, 2025 announcement confirms the event and 53,199 participants; that is a participant count, not a project count. [Cognizant announcement](https://news.cognizant.com/2025-08-21-Cognizants-Vibe-Coding-Event-Sets-GUINNESS-WORLD-RECORDS-TM-Title) The Cognizant AI Lab retrospective reports more than 30,000 submissions processed using its Neuro AI Multi-Agent Accelerator. Specialized evaluators covered innovation, UX, scalability, market potential, implementation feasibility, financial feasibility and technical complexity. Inputs were descriptions/PDFs, video-pitch transcripts and repository metadata. The system generated scores and rationales, reportedly in under 15 hours on ten servers for $7,000. Humans then reviewed top-scoring projects for final selection. **Unknown:** exact models, aggregation weights, override rates and whether repository contents were executed; metadata analysis must not be presented as runtime verification. Cost and speed are organizer claims. [AI Lab account, September 3, 2025](https://decisionai.substack.com/p/how-we-judged-the-worlds-largest) **NEBULA:FOG:SINGULARITY — an AI watched the demos and published rankings.** The organizer describes The Arbiter listening to presenter audio and viewing screens through Gemini Live. Gemini, Claude and a Groq-served evaluator independently scored demos; aggregation included outlier detection. The rubric weighted technical execution 40%, innovation 30% and demo quality 30%, with an additional track bonus up to 10%. A comparative deliberation followed. The results page names AgentRange first and publishes per-team scores. **Caveat:** it reports 23 teams but 25 demos, without a complete reconciliation, and some commentary says observable evidence is missing even where detailed positive technical judgments appear. Exact model versions, bonus arithmetic and human override authority are unclear. These are published AI results, not independent verification of the teams’ work or proof of prize payment. [Organizer’s results and system description](https://nebulafog.ai/singularity-results.html) **Synthesis — partner agents proposed; humans decided.** Devfolio’s April 30 retrospective reports 685 projects and 28 partner judges powered by Bonfires.ai. Each partner supplied its own rubric and priorities. Agents reviewed documentation; shortlisted projects received code and deployment checks, including comparisons between README claims and deployed behavior, reportedly taking 15–20 minutes per shortlisted project. Humans made the final decisions. **Unknown:** complete prompts, exact model versions, aggregate weighting and the rate at which humans rejected recommendations. The retrospective is substantially clearer about authority than the “AI judges” headline. [Devfolio retrospective](https://devfolio.co/blog/synthesis/) The official site gives a March 13–22 building window; Devfolio’s event listing instead displays February 9–March 23. These may describe different phases, but that is an **inference**, not a resolved discrepancy. A sponsor’s April 11 account independently corroborates its own judging-agent involvement and identifies projects it selected. [Official schedule](https://synthesis.md/), [event listing](https://synthesis-md.devfolio.co/), [EigenCloud sponsor account](https://www.eigenlabs.org/blog/ai-gent-hackathon-sythesis/) **HackerRank Orchestrate — defend your build to an AI interviewer.** Four editions are separately evidenced, avoiding the mistake of counting the recurring series only once: | Edition | Build dates | Evidence of occurrence and advertised judging | | --- | --- | --- | | May 2026 | May 1–2 | Retrospective describes completed judging; [official contest archive](https://www.hackerrank.com/contests) records May 2 end | | June 2026 | June 19–20 | [June event page](https://www.hackerrank.com/hackerrank-orchestrate-june26) links its leaderboard and describes the AI interview; archive records completion | | August 2026 | August 1–2 | [August event page](https://www.hackerrank.com/hackerrank-orchestrate-august26) links its leaderboard and includes winner testimony; archive records completion | | September 2026 | September 12–13 | [September event page](https://www.hackerrank.com/hackerrank-orchestrate-september26) specifies the interview and September 15 results date; archive lists the event as ended | The May retrospective describes four separately scored signals: source code, task-output CSV, coding-assistant chat transcript and a 30-minute AI voice interview. Code evaluation checked implementation; outputs were compared with a reference dataset, including adversarial cases. The interview tested architecture, failure modes, ownership and tradeoffs. Scores were combined using fixed weights, with the same scoring method across the cohort. **Unknown:** exact judge model/version, all final combination weights and human override procedures. Later event pages preserve the artifact-plus-interview format, but this report does **not** assume May’s task-specific rubric or reference dataset carried unchanged into later editions. [HackerRank’s completed May evaluation account](https://www.hackerrank.com/blog/behind-the-scenes-of-hackerrank-orchestrate/) **Cursor Boston × PyData — a separate AI notebook prize.** At the May 13, 2026 event in Cambridge, Massachusetts, entrants built Marimo notebooks using specified datasets. The organizer separated three human-judged presentation prizes from three notebook prizes selected by a blind AI agent. Submissions arrived through repository pull requests; the public gallery contains AI scores and feedback. **Unknown:** judge model, prompt, rubric weights, whether notebooks were actually executed, and exactly which identifying metadata was concealed. “Blind” is the organizer’s description, not an anonymity audit. The gallery remains open to later exercise submissions, so its present contents are not necessarily the original competition field. [Event rules and scored gallery](https://www.cursorboston.com/events/cursor-boston-pydata-2026) **Code Forward 2026 — AI deliberation supports a human judge.** Money Forward India’s judge Yosuke Suzuki describes deploying Judgie-AI at the internal event. Source ZIPs, videos and PDFs fed five expert personas: entrepreneur, engineer, UX designer, product manager and investor. They returned scores and multilingual feedback; teams could consult the system up to three times and submit one objection for reconsideration. The human judge used the output as a draft. The account mentions Gemini and publishes both favorable survey responses and reports of wrong-project feedback or incomplete code coverage. **Unknown:** exact event date, final award weighting and independent accuracy measurements. The released software was subsequently improved, so its present behavior is not proof of the event-time implementation. [Builder/judge retrospective, July 2, 2026](https://global.moneyforward-dev.jp/2026/07/02/how-a-hackathon-judge-hacked-the-hackathon/) **Devin × Claw Collective × Qwen — an AI courtroom on Discord.** The system builder reports judging over 40 cases at #BuildForMsia on August 23. Submissions supplied a problem, solution, live URL and video. The documented pipeline sanitized and deduplicated intake, smoke-tested URLs, ran three independent persona scores, aggregated them and wrote sealed scores to a sheet. Teams answered questions in private Discord threads on a seven-minute clock. Builder, Skeptic and Futurist personas examined functionality, commercial value and agent behavior. Deliberations were hidden; a voiced replay was broadcast afterward. **Unknown:** exact deployed models and rubric weights, actual final winners, and whether humans judged the final round. This establishes the builder’s deployment claim and review mechanism, not independent reproduction. [System repository and event account](https://github.com/shuenrui/hackathon-courtroom) **SANS FIND EVIL! — an individual judge’s evidence-based agent panel.** Kenneth G. Hartman’s retrospective describes Gemini as referee, Grok as prosecutor, Claude as defender and GPT as verifier, reviewing repositories and execution evidence. A separate Gemini pass summarized videos. A reconciler prioritized evidence quality over simple averaging; a ranking council proposed comparative adjustments, with close calls left to the human judge. He reports a truncated-input review missing an existing license and an agent prematurely ending after following repository text. **Unknown:** exact versions, prompts, per-entry traces and how many other judges used this workflow. [Judge’s August 28 account](https://kennethghartman.com/blog/judging-find-evil-with-an-ai-panel/) The organizer describes a wider panel of human incident responders and hands-on finalist testing. Therefore it would be misleading to call the entire competition autonomously AI-judged. [Organizer’s finalist and validation announcement](https://findevil.devpost.com/updates/45839-announcing-find-evil-hackathon-finalists) **Additional actual-use reports with weaker agent or event classification** | Case | What is supported | What remains uncertain | | --- | --- | --- | | Standard Bank, “Code. Conquer. Grow,” 2025 | The bank’s sustainability report, printed p. 71, says an AI judging app narrowed 623 ideas to 50; 15 later reached a Dragon’s Den finale. [Primary report](https://www.standardbank.com/static_file/StandardBankGroup/filedownloads/RTS/2025/SBG_SustainabilityDisclosuresReport2025.pdf) | Model, rubric and whether it was an agent rather than a scorer are undisclosed. The passage was available through indexed PDF extraction; direct opening failed. This supports AI screening, not autonomous final judging. | | IEB–TechWays National AI Hackathon, May 20, 2026 | The [organizer](https://techways.online/ai-hackathon/) confirms occurrence. [Melio AI’s own deployment statement](https://www.linkedin.com/posts/melio-consulting_were-proud-to-see-melio-ai-featured-in-iafricacom-activity-7475087830386573312-ozSh) says its joint Project RubrIQ platform supported judging of 200+ team submissions and assisted human experts. | Agent architecture and exact aggregation were not established in primary technical documentation. A [June 18 results article](https://iafrica.com/ai-joins-the-judging-panel-for-2026-ieb-techways-ai-hackathon-2026-winners-announced/) provides a more detailed five-judge/shortlist account; treat that as a secondary lead for confirmation, not audited implementation evidence. | | SVCAF AI4Legislation 2025 | The [organizer’s award account](https://svcaf.org/posts/svcaf-celebrates-innovation-at-ai4legislation-2025-award-ceremony/) confirms human and AI evaluation. Its [ceremony slides, slide 4](https://svcaf.org/files/github-svcaf_ai4legislation_award_ceremony.pdf) identify three AI judges. The rubric was innovation 25%, impact 25%, technical excellence 20%, usability 15%, ethics 15%. | Described as a competition, not clearly a time-bounded hackathon. AI validated results; its binding score contribution and agent architecture are not specified. Award ceremony: September 21, 2025. | | FutureMinds Hackathon | Supplier Veav AI reports AI Scott and AI Maria scored against predefined criteria, gave real-time feedback and mentored entrants alongside humans. [Supplier case study](https://veavai.com/case_study/hackathon/) | Undated, lightly documented supplier account; event identity, venue, winners, models and aggregation missing. Direct opening returned a redirect page; details were available in indexed content. Fairness claims are promotional, not demonstrated. | **Published judging designs: do not yet count as completed AI judging** **FFAI Robothon Summer 2026.** This is a particularly explicit all-AI design. Its official repository says GPT, Claude and Gemini score independently with no human scoring. Entrants submit runnable MuJoCo simulations, code, run instructions and videos by pull request. [Official repository](https://github.com/Faraday-Future-AI/Robothon-starter) The event page specifies identical review packages, three calls per judge, each judge’s median, and then the mean of those three medians. Eight dimensions cover runnability, MuJoCo depth, task design, control, dexterity, engineering, presentation and innovation. It advertises fixed temperature, submission limits and latest-score-only ranking. **Unresolved:** the retrieved leaderboard showed only a loading placeholder; no completed score trail or final results were established. June 16 submission opening is not evidence of completed judging. The page was indexed but direct retrieval failed. Repeated scoring is a stated variance-control mechanism, not proof of unbiased or deterministic outcomes. [Official event description](https://robothon.ff.com/) **Jom Quick & Kiro TechJam, August 27, 2026, Putrajaya.** The event guide specifies screenshot-and-description submissions to an AI judge, live scores and iterative resubmission, then a top-three final with presentations and Q&A. Four equal rubric categories cover innovation, UX/design, tool usage and impact, with an integration bonus. **Unresolved:** no completed results or judging logs located; model, agent architecture and final judge identity are unspecified. Treat this as a published two-stage design despite the elapsed event date. [Event guide](https://main.d2h13mvts4287z.amplifyapp.com/) **Orion Builder Hackathon, August–September 2026.** The organizer specifies an AI vetting score at submission, community upvotes and partner judges who choose winners. Its written deadline is September 27; the retrieved page still displayed a countdown, so page freshness is uncertain. **Unresolved:** completed results, evaluator architecture and the AI score’s weight. This is specified AI-assisted judging, not an AI-selected winner. [Organizer’s rules](https://orionagents.org/hackathon) **CHIA / A³ at MICRO 2026.** Published rules describe AI agents with different personas providing reviews to human program-committee members. The rules explicitly exclude AI judgments on acceptance, rejection or ranking. **Classification:** advisory review design, not an autonomous AI jury; completed use was not established. [Official judging process](https://agentic-arch.org/hackathon.html) **Experimental rounds and judging products** **Overhacked** (Devpost title spelled “Ovehacked”) is worth retaining as an actual miniature agent competition, rather than silently equating it with a public human-entry event. Its builders report a round of ten agents in three teams that shipped and were judged in under eight minutes. Each agent used a Daytona sandbox; live-URL health checks gated entry. Product, craft and engineering jurors scored independently, with highest total winning and evaluations logged in Braintrust. **Uncertain:** exact round date, model versions and external prize/event status. The report of a working round is first-party and was not replayed here. [Builder’s project account](https://devpost.com/software/ovehacked) These other discoveries establish **tools or prototypes**, not completed qualifying hackathons: | Product | Evidence and reason not counted | | --- | --- | | AIJudge, ETHGlobal Taipei | Its [project showcase](https://ethglobal.com/showcase/aijudge-oeihx) describes automated review and onchain rewards. Being built at Taipei does not show it judged Taipei. | | Agent Arena | [Devpost project](https://devpost.com/software/agent-arena-dv3sr1) describes multi-agent repository evaluation; no specific completed event established. Do not conflate it with unrelated events bearing the same name. | | HackSight | [Builder account](https://devpost.com/software/hacksight) describes code, market, chat and score agents; no named deployment established. | | ORCHESTRA | [Builder portfolio](https://www.yashvasudeva.app/projects/orchestra) describes five judges, a bias auditor and chief judge; no completed event identified. | | JuriXAI | [Builder announcement](https://www.linkedin.com/posts/kane-pascal-041459208_jurixai-autonomous-hackathon-judging-activity-7496601227053568000-q-Kv) describes four specialized judges and escrow payouts; no completed hackathon result verified. | | AI Jury / EvalLens | The [March builder article](https://builder.aws.com/content/3B2R0iO8Mri8BNthbttLKY6JWyy/how-we-built-ai-jury-a-multi-agent-pitch-evaluation-system-on-amazon-nova-2-lite) described pilot discussions. The [June retrospective](https://www.evallens.io/blog/from-ai-jury-to-evallense) describes internal evaluation runs; neither establishes a named completed external event. | | KIT Hackathon platform | Its [help page](https://www.kitcbehackathon.in/help) mentions an automated AI jury, but a specific completed AI-judged edition was not established. | An unnamed workplace competition described by [Ted Ward](https://tedward.net/blog/breaking-the-bench-ethical-dilemmas-and-vulnerabilities-in-ai-as-a-judge-systems) remains an unverified lead: the accessible indexed account reports an announced AI judge but does not identify the employer/event or prove final judging occurred. A research paper, [AI Judges: A System with Absolute & Relative Evaluation in Ideathons and Hackathons](https://doi.org/10.1109/ACDSA67686.2026.11468180), reports evaluating 8,942 ideathon submissions; the accessible abstract does not establish a named hackathon deployment. Neither is added to the core count. **Important exclusions.** Colosseum’s Solana Agent Hackathon had agent builders and votes, but its [official announcement](https://blog.colosseum.com/announcing-colosseums-agent-hackathon/) says judges select winners without identifying them as AI. Microsoft’s [AI Agents Hackathon 2025](https://microsoft.github.io/AI_Agents_Hackathon/) identifies human engineers, managers and advocates as judges. Neither title proves AI judging. Art contests such as [AI-ARTS](https://ai-arts.org/competitions/4th-ai-arts-competition-2025/winners) and continuous agent contests such as [Omniology](https://www.omniology.ai/home) are adjacent competitions, not established hackathons under this scope. **What the evidence suggests — analysis, not additional facts.** - **Human authority matters more than the headline.** The documented handoffs range from screening, to draft scorecards, to separate AI prize tracks. A meaningful comparison must identify which decision the AI could actually change. - **Evidence access is central.** A judge reading a description, one inspecting a repository, and one questioning an entrant are not performing equivalent evaluations. The cases above support distinguishing these inputs; they do not establish that any one method always produces better rankings. - **A panel can expose disagreement without proving correctness.** Independent scores, persona debate and evidence-weighted reconciliation are different mechanisms. More agents alone does not establish fairness, independence or accuracy. - **Feedback can shape the entry before judging ends.** Mid-event coaching and resubmission make the judge part of the development process. This may improve work, but can also favor optimization for the published evaluator; that is a design implication, not a measured outcome here. **Unanswered questions for a definitive follow-up.** The most useful missing records are event-time prompts and model versions; complete per-entry scores and evidence; actual execution logs; human overrides and appeals; participant consent/disclosure; and final results linked to award decisions. For FFAI and TechJam, completed results are especially important. For Standard Bank, FutureMinds and IEB–TechWays, implementation details would establish whether “AI judge” meant an autonomous agent, a panel of model calls, or a simpler scoring application. **Research limits.** This report reconciles public accounts, not private logs. Organizer and vendor claims remain attributed even when detailed. Dynamic pages may contain later edits; crawl labels are not event dates. No judging code was executed, competitors contacted, accounts created, or prizes verified. See the accompanying README for the search method and checks. The report’s structural validation does not certify the truth of its sources.