peakstateglobal Work with us

We are legislating the wrong half of AI disclosure

Almost every argument about disclosing AI is two arguments wearing one name. Separate them and you get a policy worth having. Leave them fused and you get a sticker.

Earlier this year I spent a couple of weeks inside the AI disclosure rules, building a workshop on how to make AI-assisted work defensible. I went in expecting the regimes to be converging on something useful, and what I found genuinely shocked me. The law is converging on a sticker.

The EU AI Act requires synthetic content to be marked as artificially generated or manipulated. Yes or no, machine or not. Meanwhile Elsevier, publishing research, asks an author using an AI-generated image to name the model, the version and the developer, and the medical journal editors ask which tools were used and how. Publishing asks how the work was made. The law asks whether a machine touched it. Two serious institutions took the same problem in the same year and went opposite ways, and the binary version is winning, because it is the easy one to legislate.

That divergence exists because "AI disclosure" is two questions wearing one name.

The first is a capability question: which model, which tools, which version, and what did the human check. This is what a reader needs to judge how far to trust the work, and its answer changes with every model release.

The second is a provenance question: whose work was used to build the model. This is what an artist or a novelist is asking, it has nothing to do with output quality, and its answer never changes, because it is a fact about how a training corpus was assembled.

The tick box answers neither. It works like the label on a supermarket shelf: "Made in Australia" means only that the last substantial transformation happened here, while the shopper reads it as a promise about the ingredients. Nobody lied. The label answers a narrow question about one step, and gets read as a guarantee about the whole product.

The tick box is doing damage while it sits there

A badge feels like the cautious option, the small step that cannot hurt. The measurements say otherwise, three ways.

It taxes the wrong work. Across sixteen preregistered experiments with 27,491 people, identical creative writing scored lower the moment readers believed AI was involved. Same words, worse rating.

It certifies everything unbadged. When researchers put warnings on some false headlines, sharing of the unlabelled false headlines rose from 29.8 per cent to 36.2 per cent. Labelling part of the pile made people trust the rest of the pile more, and every labelling scheme ever run has been partial.

And it misses its actual target. A study of 800 participants found an AI-content label had no significant effect on whether people believed or shared misinformation. The badge moves judgement where the value is authorship, and does nothing where the value is truth.

Is the badge at least a rough guide to quality?

Nope.

Hallucination rates across 26 current models range from 22 per cent to 94 per cent. "AI was used" now covers everything from unusable to excellent, and the range moves every quarter. On SWE-bench Verified, a software engineering benchmark, performance went from roughly 60 per cent to close to 100 per cent in a single year. Naming the specific model tells a reader something they can look up. Naming the category tells them which prejudice to apply.

The direction is one-way, because checking an answer is cheaper than producing one. On a graduate-level science benchmark, generating many candidate answers and filtering them with weak automated judges scored 82.8 per cent where majority voting scored 45.5. Wherever answers can be checked cheaply, reliable systems get built out of unreliable parts. The United States cleared a fully autonomous diagnostic system in 2018, no clinician reading each case, and in the MASAI trial an AI-first screening arm beat standard double reading by two human radiologists across roughly 100,000 women. I expect an Australian regulator to permit AI work product without per-item human sign-off before the end of 2028, which means the tick box you are drafting this quarter will outlive the assumption underneath it.

But surely a machine decision needs more justification than a human one?

This assumption is the load-bearing wall under the whole checkbox, and it does not survive contact with the evidence.

Watch what the two answers trigger. Tick "yes, AI was used" and you invite scrutiny, review steps, a second pair of eyes. Tick "no" and nobody asks anything at all about how the judgement was reached.

Human judgement does not earn that free pass. When a machine learning conference sent its own submissions to two independent review committees, 50.6 per cent of the papers the first committee accepted were rejected by the second. The outcome turned on who happened to read the paper, in a field that builds measurement systems for a living.

And the explanation a decision-maker gives you was built after the decision. In the choice-blindness experiments, people chose which of two faces they preferred, the card was covertly swapped for the one they rejected, and no more than 26 per cent ever noticed. The rest explained, fluently and sincerely, a choice they had never made. The neuroscience points the same way: a choice can be read out of brain activity up to ten seconds before it reaches awareness. Deliberate-then-decide is not the only ordering available to the brain, and given the millions of decisions we make on autopilot every day, it is not even the most common one. The decision arrives, and then we get our turn to explain it. We use frameworks, due process and panels of peers precisely because we know this about ourselves; that is what those structures are for, reducing the bias and nepotism the individual account will never confess to.

Language models work the same way, and here the parallel is exact. A model streams its reasoning and will happily justify its answer after the fact, just as we do. When researchers slipped a biasing hint into a question, models changed their answers to match the hint and then wrote reasoning that justified the new answer without mentioning the hint. Anthropic measured it directly on models built to reason out loud: Claude 3.7 Sonnet acknowledged the hint it had used 25 per cent of the time, DeepSeek R1 39 per cent. The unfaithful explanations were longer than the faithful ones.

So the reasoning you get from a person and the reasoning you get from a model are the same kind of object: a justification composed after the answer. I include myself in this. The change it forced on me was to stop treating my own explanation as evidence and to go looking for the thing that would prove me wrong instead.

That is what collapses the tick box. It sorts work into a column that owes you an explanation and a column that gets waved through, and neither column can produce an explanation you could check.

Ask for the ingredients list

The fix is on the same supermarket shelf as the problem. An ingredients list names what went into the product and lets anyone who cares read it, look up the additive they distrust and decide for themselves. That is the disclosure worth having for AI-assisted work: the model, the tools and the versions, what was checked against what, and by whom. Every element is something a reader can interrogate. Look up the model's known failure modes, open the cited source, rerun the number.

A badge gives the reader nothing to do except apply a prejudice, and that is its deepest cost. It invites people to outsource their judgement to a stamp instead of reading the work. An ingredients list pulls the other way: it hands the reader the material to find what might invalidate the work, and trusts them to use it.

Make it reachable rather than staged at the front. A banner at the top sets the reader's prejudice before the first word, which is the disclosure penalty manufactured on purpose. A line at the end, or a link, lets them assess the work first and interrogate the pipeline second.

And send the provenance question upstream, where it can actually be answered. The person at the keyboard cannot tell you what the model was trained on. Ask the vendor, in procurement: what was this trained on, and what will you indemnify? The courts are still sorting the rest out. Training on lawfully acquired books has been held transformative fair use while pirated libraries were not, the leading settlement is approved but on appeal, and no binding appellate ruling exists anywhere yet. A contract clause protects you today; the case law will take years.

Will our lawyers sign this?

Their objection deserves stating plainly: a detailed disclosure is a map. Name your models, tools and checks, and you have handed a future plaintiff the discovery roadmap, where a bare badge can be defended as "we disclosed". That exposure is real, and where it lands for your organisation is a conversation with your counsel, not with me.

Two things belong in that conversation. Minimal disclosure is already a closing window: APRA told industry in April 2026 that governance is not keeping pace and embedded AI is reducing transparency, the National AI Centre's essential practices ask you to name who is accountable, and from 10 December 2026 the Privacy Act requires your privacy policy to state which kinds of decisions are made solely by a computer program. And the same objection was raised against every reporting regime we now treat as ordinary. It lost each time on the same ground: the organisations that could not explain their own work fared worst when somebody finally asked.

What would make me wrong

Everything above about machine reliability rests on answers being cheap to check, and nobody has measured how much real knowledge work is checkable. If most of what your organisation does is long-horizon and open-ended, the capability signal stays meaningful longer than I have claimed. That measurement does not exist yet, and I would change my position if it came back against me.

Where to start

Stop asking your people whether a machine was involved. Ask what they checked. That question has an answer regardless of what produced the work, and the answer is worth reading.

The SOURCED prompts and templates are free, no signup. They make a document carry its own ingredients list: what is sourced, what is recalled, what is inference, and where the limits are. Take them, adapt them, or write your own. The tool matters less than dropping the box.


text
Attribution:  Written by Andrew Ramsden, Peak State Global. AI tools assisted the
              research and drafting; all claims were checked against the retrieved
              sources named below. Pipeline: Claude Opus 5 (claude-opus-5, 1M
              context) running in Claude Code (VSCode extension, no build version
              exposed), with the SOURCED research and audit prompts (sourced repo,
              commit 29728f7) and the retrieval-ladder reference (no version).
Accountable:  Andrew Ramsden, andrew@peakstate.global.
Limitations:  The MASAI interval-cancer figures were read through a PubMed Central
              commentary quoting the trial, not the Lancet paper itself, which was
              paywalled at the time of retrieval. The Bartz, Kadrey, Ross and Getty
              rulings were read through law-firm analyses quoting the judgments,
              not the judgments themselves, except for the Bartz docket entries,
              which were read directly, and the Getty analysis through a web
              archive copy because the publisher blocks automated retrieval. The
              2 December 2026 extension to the EU marking duty is asserted by a
              secondary source without an Official Journal citation. The 2028
              expectation about an Australian regulator is a forecast; its date
              and settling criterion are recorded in the claim sidecar.
References:   Twenty-two sources, retrieved on 30 August 2026, each figure quoted with
              its locator in the claim sidecar, which lists every claim in this piece
              against the evidence it rests on:
              [SOURCED sidecar](https://github.com/peakstate-global/sourced/blob/main/docs/examples/wrong-half-ai-disclosure.sourced). Court dates and settlement status
              were confirmed against the CourtListener docket for Bartz v. Anthropic,
              4:24-cv-05417 (N.D. Cal.), through 25 August 2026.

              AI and breast cancer screening at a crossroads: Insights from the
                MASAI trial. (2026). PubMed Central.
                https://pmc.ncbi.nlm.nih.gov/articles/PMC13036691/
              Anthropic. (2025). Reasoning models don't always say what they think.
                https://www.anthropic.com/research/reasoning-models-dont-say-think
              Australian Competition & Consumer Commission. (n.d.). Country of
                origin claims. Retrieved August 30, 2026, from
                https://www.accc.gov.au/business/advertising-and-promotions/country-of-origin-claims
              Beygelzimer, A., Dauphin, Y., Liang, P., & Wortman Vaughan, J. (2021,
                December 8). The NeurIPS 2021 consistency experiment. NeurIPS Blog.
                https://blog.neurips.cc/2021/12/08/the-neurips-2021-consistency-experiment/
              Coalition for Content Provenance and Authenticity. (2026). C2PA
                specification (Version 2.3).
                https://spec.c2pa.org/specifications/specifications/2.3/specs/C2PA_Specification.html
              Digital Diagnostics. (n.d.). FDA permits marketing of LumineticsCore
                (formerly known as IDx-DR) for automated detection of diabetic
                retinopathy in primary care [Press release].
                https://www.digitaldiagnostics.com/fda-permits-marketing-of-lumineticscore-formerly-known-as-idx-dr-for-automated-detection-of-diabetic-retinopathy-in-primary-care/
              DLA Piper. (2025). Getty Images v Stability AI: The UK High Court
                decision offers guidance, but critical questions on AI and copyright
                infringement remain.
                http://web.archive.org/web/2/https://www.dlapiper.com/en-us/insights/publications/2025/11/getty-images-v-stability-ai-the-uk-high-court-decision-offers-guidance-but-critical-questions-on-ai
              Elsevier. (n.d.). Generative AI policies for journals.
                https://www.elsevier.com/about/policies-and-standards/generative-ai-policies-for-journals
              Goodwin Law. (2025). Northern District of California judge rules that
                Meta's training of AI models is fair use.
                https://www.goodwinlaw.com/en/insights/publications/2025/06/alerts-practices-aiml-northern-district-of-california-judge-rules
              Impact of artificial intelligence-generated content labels on perceived
                accuracy, message credibility, and sharing intentions for
                misinformation. (2025). PubMed Central.
                https://pmc.ncbi.nlm.nih.gov/articles/PMC11892328/
              International Committee of Medical Journal Editors. (n.d.). AI use by
                authors.
                https://www.icmje.org/recommendations/browse/artificial-intelligence/ai-use-by-authors.html
              International Press Telecommunications Council. (n.d.). Digital source
                type NewsCodes. https://cv.iptc.org/newscodes/digitalsourcetype/
              Johansson, P., Hall, L., Sikström, S., & Olsson, A. (2005). Failure to
                detect mismatches between intention and outcome in a simple decision
                task. Science, 310(5745), 116–119.
                https://doi.org/10.1126/science.1111709
              LawSites. (2026, June). At 3rd Circuit, judges press ROSS and Thomson
                Reuters on fair use, AI training and market harm.
                https://www.lawnext.com/2026/06/at-3rd-circuit-judges-press-ross-and-thomson-reuters-on-fair-use-ai-training-and-market-harm.html
              Pennycook, G., Bear, A., Collins, E. T., & Rand, D. G. (2020). The
                implied truth effect: Attaching warnings to a subset of fake news
                headlines increases perceived accuracy of headlines without warnings.
                Management Science, 66(11), 4944–4957.
                https://doi.org/10.1287/mnsc.2019.3478
              Raj, M., Berg, J. M., & Seamans, R. (2026). The artificial intelligence
                disclosure penalty: Humans persistently devalue AI-generated creative
                writing. Journal of Experimental Psychology: General, 155(4),
                896–915. https://doi.org/10.1037/xge0001889
              Regulation (EU) 2024/1689 of the European Parliament and of the Council
                of 13 June 2024 laying down harmonised rules on artificial
                intelligence (Artificial Intelligence Act). (2024). Official Journal
                of the European Union, L 2024/1689.
                http://web.archive.org/web/20260822045623/https://eur-lex.europa.eu/legal-content/EN/TXT/HTML/?uri=CELEX:32024R1689
              Soon, C. S., Brass, M., Heinze, H.-J., & Haynes, J.-D. (2008). Unconscious
                determinants of free decisions in the human brain. Nature Neuroscience,
                11(5), 543-545. https://doi.org/10.1038/nn.2112
              Stanford Institute for Human-Centered Artificial Intelligence. (2026).
                Artificial intelligence index report 2026.
                https://hai.stanford.edu/assets/files/ai_index_report_2026.pdf
              Turpin, M., Michael, J., Perez, E., & Bowman, S. R. (2023). Language models
                don't always say what they think: Unfaithful explanations in
                chain-of-thought prompting. arXiv. https://arxiv.org/abs/2305.04388
              Weaver: Shrinking the generation-verification gap with weak
                verifiers. (2025, June 18). Hazy Research, Stanford University.
                https://hazyresearch.stanford.edu/blog/2025-06-18-weaver
              Wiggin and Dana. (2025). Bartz v. Anthropic: First court decision on
                fair use defense in LLM training.
                https://www.wiggin.com/publication/bartz-v-anthropic-first-court-decision-on-fair-use-defense-in-llm-training/
published