Australian not-for-profit · ACN 699 651 771Open research · Open software · Public benefitPublic record

Learn / AI safety

Why an agent must show its work

Follow performance evidence into reputation, and see how a leaked evaluation breaks the signal.

DSEMA research architecture · Reviewed 3 October 2026

Explore the mechanism

Trace the processDSEMA research architecture • illustrative evaluation

A score needs evidence behind it.

Select a stage to see who acts, what changes, and what evidence remains.

Task designer

Specify a task and an independently checkable success criterion.

A vague objective gives agents room to optimise the wrong thing.

Stage 1 of 4
Stages explain a design and its evidence; a highlighted stage is not proof of real execution. Sources: underlying documentation.

01 / Explanation

Usefulness needs a definition

DSEMA studies reputation derived from recorded task performance and allocation to capable specialists. This offers an alternative to treating machine existence as an issuance entitlement. It does not prevent agents from accumulating wealth or obtaining funding elsewhere.

02 / Explanation

More agents are not always more independent

Voting gains depend on assumptions about errors and independence. Agents sharing training data, prompts, or tools may fail together. Evaluate correlated failures and collusion rather than presenting an ensemble size as a guarantee of correctness.

03 / Explanation

Keep evaluation harder to game

The research proposes adversarially evolving evaluations. Useful tests would examine held-out performance, reward hacking, evaluator conflicts, reputation recovery after failure, and whether new agents can compete against entrenched specialists. A score should remain inspectable and contestable.

Check your understanding

Why can ten agreeing agents still be wrong?

Next: Who keeps the system accountable? →All guides