Orchestrator Awards Working Papers · Current Orchestrator Institute record · About the Author

Judging AI-Native Work

Gates, Evidence, and the Limits of the Machine

Working Paper v0.1

Issue: Issue 003

Version DOI: 10.5281/zenodo.22049566

Review status: Working paper — not peer-reviewed.

Current research-library entity: Orchestrator Institute

Published by: Orchestrator Institute Research Labs

Author: Mark Sendo · ORCID 0009-0001-6724-307X

Affiliation: Celestial Technologies LLC

Date: August 21, 2026

Publication track: Research / field paper

License: CC BY 4.0

Suggested citation: Sendo, M. (2026). Judging AI-Native Work: Gates, Evidence, and the Limits of the Machine. Orchestrator Awards Working Paper, Issue 003, v0.1. OrchestratorAwards.com. https://doi.org/10.5281/zenodo.22049566

Disclosure: This issue is published by the Orchestrator Institute. The author holds an institutional interest, stated so the reader can weigh it. The text was drafted with AI assistance; all claims, positions, and errors are the author’s.

Institutional note: The Orchestrator Institute and Orchestrator Awards are sibling institutions under Celestial Technologies LLC. Publication does not activate scoring, evaluation, judging, submissions, payments, nominations, juries, voting, or recognition.

Doctrine: Publish first. Be cited second. Recognize third.


Abstract

A new category of creative and technical work has arrived faster than the means to judge it. Work whose defining technique is machine generation cannot be credibly assessed by criticism that has never had to measure such technique, and it cannot be handed to a model to assess without letting the thing under evaluation become its own judge. This issue argues that credible recognition of AI-native work rests on three commitments that must hold together: admissibility is decided before merit, evidence is required before any claim is credited, and authority over recognition remains permanently human. The first two are procedural; the third is a matter of custody. Extending the argument of The Custody of Intelligence, we hold that recognition authority is a delegable power that must not be delegated to the systems being recognized, regardless of how capable those systems become.

I. The judgeability problem

An award is not a reaction, and it is not a matter of taste. It is a determination — about who is accountable for a piece of work, and whether that work meets a defined standard — made by qualified people, against published criteria, on the basis of evidence, with a record of how the decision was reached. If the reasoning behind a verdict cannot be shown, the verdict is not credible. This is the premise the rest of the argument depends on.

AI-native work strains conventional evaluation along two independent lines. The first is competence. Assessing a serious AI-native work requires reading things that traditional criticism has never had to measure: the orchestration behind the result, the workflow that held identity or continuity together, the difference between a method that genuinely solved a hard problem and one that got lucky in a single instance. That literacy has, so far, accrued mostly among the people building the work. The second is accountability. When a system contributes materially to what is produced, the question of who answers for the result does not disappear — it becomes central. Traditional evaluation can assume a human author; here that assumption must be established, not presumed.

POSITION. These are not two problems but one, and it is the custody problem in the specific setting of recognition: authority over a judgment must rest with an identifiable, competent, accountable human, or the judgment is not one the institution will stand behind.

II. Admissibility before merit

Merit is the wrong place to begin. Before any work is scored, it must clear five gates, and each is pass/fail: a named person or disclosed team accepts responsibility for the work; all material AI use is disclosed completely and truthfully; rights, likeness, voice, and consent are cleared and attested; the minimum workflow evidence is present and version-locked; and the work satisfies its category’s specific eligibility statement. Failing any one gate ends evaluation. Scoring begins only after all five pass.

The ordering is deliberate. Gates decide whether a work may be evaluated at all; scoring measures how well an admitted work performs. Collapsing the two would let craft launder a failure of accountability — a beautiful result excusing an undisclosed method or an uncleared right. Keeping them separate means the floor is a floor: no quality of output purchases an exemption from it.

III. Evidence over claim

A work is judged on what can be shown, not on what is asserted. Entrants supply the record — workflow logs, tool disclosures, provenance, and the materials each category requires — that lets an evaluator assess the craft rather than take the result on faith. Where an achievement cannot be evidenced, it cannot be scored. This is not suspicion of entrants; it is the same discipline the institution applies to its own verdicts. A recognition that rests on unverifiable claims is no more credible than a claim that rests on none.

This mirrors a principle from The Custody of Intelligence: different layers of a system are held to different standards of evidence. Evaluation inherits that logic. The more consequential the recognition, the heavier the evidentiary burden it must carry.

IV. The measurement architecture

Once a work is admitted, it is scored across ten weighted pillars: Orchestration Craft (16%), Story and Intent (14%), Identity and Continuity (11%), Technical Execution (10%), Transparency and Disclosure (10%), Human Creative Control (10%), Originality and Creative Novelty (9%), Safety and Rights Discipline (8%), Workflow Evidence (8%), and Field Contribution (4%). The weights are published so the instrument can be examined and argued with, not merely trusted.

One pillar deserves particular note because it also appears as a gate. Disclosure is tested twice, and the two tests measure different things without double-counting. As a gate, disclosure is about honesty: was material AI use disclosed completely and truthfully? Incomplete or untruthful disclosure is disqualifying, whatever the quality of the work. As the Transparency and Disclosure pillar, it is about quality: once disclosure has been given, how precise, versioned, and auditable is it? In a form where the method is part of the art, concealing the method is itself a failure of the work — so honesty is a condition of entry and clarity is a dimension of excellence.

V. The limits of the machine

Here the paper makes its central claim. The most consequential rule in evaluating AI-native work is not about how the work is scored but about who may score it.

Final judgment is always the act of named, unconflicted human evaluators. No model decides eligibility, nomination, finalist status, or recognition. If model assistance is ever introduced, it may only help organize information — surface an entrant’s own disclosures, arrange evidence for human review — and never exercise judgment. Its scope, evidence handling, auditability, and limits would be published and separately authorized before any use. The firewall between assistance and authority does not move.

POSITION. Recognition authority belongs to the class of powers that must remain human regardless of capability. The Custody of Intelligence identifies thresholds of authority that may not be autonomously delegated no matter how capable a system becomes; the authority to confer recognition on AI-native work is one of them. The reason is structural, not sentimental. To let a model adjudicate the merit of model-generated work is to let the field grade its own examination — to collapse the very distance that makes a verdict meaningful. Capability cannot resolve this, because the objection is not that models judge poorly but that their judging, however good, forecloses accountability. A verdict must be answerable to someone; a model cannot be that someone.

Several further disciplines follow from the firewall and keep human judgment honest. Evaluators disclose conflicts and confirm confidentiality before reviewing, and recuse where a conflict exists, including any relationship to the institution’s ownership; the recusal record is part of the audit trail. Evaluators score independently, unable to see or influence one another before their assessments are locked, and individual scores are then aggregated so that no single voice determines an outcome. And because recognition is a governed act, it is open to challenge: there is a defined process to raise evidence disputes, flag errors, and correct the record. A verdict that cannot withstand scrutiny is not one the institution wishes to stand behind.

VI. Domain literacy and its limits

Competence is real but bounded. Judging AI-native work requires literacy specific to the domain the work inhabits — the ability to distinguish a workflow that solved a hard problem from one that succeeded once by chance. Evaluators are chosen for demonstrated competence in the relevant domain, and literacy in one domain is not treated as authority in another. This is why the institution develops domain-specific methodologies rather than a single universal judge: a framework built for one form is published as specific to that form and is not silently generalized to others. The measurement instrument is shared machinery; the expertise that reads it is not transferable across domains.

VII. The honorary exception

Not every honor runs through the competitive rubric, and the system is clearer for saying so. The Pioneer Award is a board-conferred honor that does not pass through the five gates or the ten pillars. It is governed by a separate published honorary methodology, with its own conflict record, recusal record, and evidence dossier. Naming the exception is itself a discipline: it marks the boundary of the competitive system rather than letting the exception erode the rule.

Conclusion

The commitments described here — admissibility before merit, evidence before claim, and permanent human authority over recognition — are published before they are exercised, and that order is the point. A methodology disclosed in advance can be examined, argued with, and cited; a verdict produced under it can be trusted precisely because the reasoning was fixed before the result was known. The deepest of the three commitments is the last. Judging AI-native work well will require better instruments and deeper literacy over time, but it will never require surrendering the judgment itself to the systems under judgment. Whatever administrative methods evolve, the authority to recognize remains a human act.

This public methodology issue does not activate scoring, evaluation, judging, submissions, payments, nominations, juries, voting, or recognition. Nothing herein is conferred as an award unless and until separately authorized.