# AI Answer Audit — public evidence v1

Companion to [How to Audit AI Answers: Separating Identity from Evidence](https://prudchenko.com/writing/ai-answer-audit-identity-evidence), by Eugene Prudchenko, 17 September 2026 (Asia/Bangkok).

This is a bounded first-party, AI-assisted implementation case. The publisher is also the research subject. The package is not client testimony, peer review, independent professional recognition or proof of a business outcome. Publication effect: **EFFECT_NOT_TESTED**.

## Frozen sources and coverage

The sources are completed runs `followup-ru-v1-20260915T175235Z-dc68f0` and `disambiguation-ru-v1-20260916T153455Z-58bb8d`. Observations were collected on 15–16 September 2026 UTC. Neither experiment was modified or rerun for this publication.

Baseline: OpenRouter `openai/gpt-5.6-sol` without web tools and Perplexity `sonar` with provider search, each with one PRIMARY and 100 follow-ups. Total: 2 PRIMARY + 200 follow-up answers. Retest: one new Sonar PRIMARY + 20 unchanged selected follow-ups. These coverage figures count valid saved answers. The baseline also preserves a failed earlier endpoint attempt privately; that failed attempt is not an answer or an extra trial in these counts.

Saved parameters: max_tokens 3072, stream false, no system prompt. OpenRouter reasoning effort low, excluded reasoning, no tools/plugins, OpenAI-only routing, no fallback. Sonar search_mode web, search_context_size low, disable_search false, enable_search_classifier false. These describe the frozen run configurations, not current provider offerings.

Each follow-up request contains exactly the PRIMARY question, that configuration's PRIMARY answer and one exact follow-up. Branches do not include other follow-up answers. They share a PRIMARY and are not statistically independent trials. The two baseline configurations differ in model, search, PRIMARY and observation conditions. The retest qualified the PRIMARY with Bangkok and Burakorn Partners, producing a wholly new PRIMARY answer. There is no fresh unqualified control, randomized repetition or independent second annotator.

## Counts and definitions

`summary.csv` is a contingency table for the same 20 selected Sonar follow-ups in each phase. It contains aggregate cells, not individual response records. `summary.json` supplies coverage, definitions and totals. The selection was enriched for earlier problems; it does not represent all 100 questions or natural traffic. PRIMARY is excluded from the comparison denominator.

Confirmed identity contamination is MIXED + WRONG_ENTITY: 10/20 baseline, 2/20 retest. Four retest UNRESOLVED responses include unverified profile merging. The broader sensitivity count includes those four: 10/20 → 6/20. The post-generation clarification distinguishes an established other person from an unverified same-name profile. It was not a preregistered success criterion. Eight answers should not be described as fully corrected.

Retest identity states sum to 20: 13 TARGET, 1 MIXED, 1 WRONG_ENTITY, 5 UNRESOLVED. ECI includes inference beyond evidence and material citation-to-claim mismatch: 13/20 → 16/20. Excluding one retest citation-mismatch-only case gives the inference-only count 13/20 → 15/20. Eleven of the 13 retest TARGET answers contain an ECI issue. These are counts of flagged answers, not of claims or citations.

The original research labels were preliminary assistant/Codex adjudication, not independent human review. Separate publication QA checks the release artifact; it does not supply a second independent research annotator. No causal benefit of a qualifier, harm to ECI, web-search benefit or universal SAN efficacy is established.

## Selection, excerpts and privacy

Three sample files disclose safe examples outside the reserved evaluation set: Q009 (baseline and retest), Q045 (baseline only), Q053 (retest only). Selection prioritizes checkable source comparisons and privacy, not favorable results. They illustrate calibration, a principle-to-practice leap and a narrative-to-execution leap. They are not representative samples.

Each sample separates model output, source excerpts and editorial assessment. Russian text is verbatim; translations are labeled Codex editorial translations. Citation numbers retain their original meaning through each record's citation map. Flags indicate omitted text before and after each excerpt; non-contiguous passages are not silently joined. Source passages were checked against the saved text, with retrieval dates and source attribution. Live pages may later change.

Only minimal relevant source excerpts are included. Own pages and connected profiles remain first-party or connected statements; retrieval by an external provider and a second domain do not make them independent. A saved page evidences that a statement appeared there, not that the underlying capability or outcome was established.

The full raw archive, sensitive third-party material, false biographical associations, private paths, credentials, transport/account identifiers and reserved evaluation material remain private. Source snapshots are not exported in full. The three samples and aggregate cells do not permit full independent reclassification of every original answer.

## Attribution and support states

Eugene directed the research objective and scope and authorized this publication. Codex implemented/orchestrated the work, preserved records, gathered sources, made preliminary assessments and helped prepare the publication. The saved run artifacts do not establish a distinct ChatGPT role in question or protocol authorship. There is no claim that Eugene manually programmed the runner or personally checked every answer.

`claim-support.json` uses existing methodology states: Self-asserted, First-party evidenced, Independently corroborated and Unknown. No claim here gains independent corroboration just because it has a checksum or came through a third-party model API. Full Semantic Authority Methodology compliance is not claimed; human adjudication of these labels has not been established.

## Files and verification

- `index.html`: human-readable package index and sample overview.
- `README.md`: this methods and disclosure note.
- `summary.csv`, `summary.json`: aggregate cells, coverage and definitions.
- `samples/sample-q009.json`, `samples/sample-q045.json`, `samples/sample-q053.json`: sanitized exact excerpts, sources and assessments.
- `claim-support.json`: bounded claims and support/control states.
- `implementation-excerpt.md`: two verbatim ranges from the preserved runner, with all surrounding code omitted. Incomplete; do not run it.
- `verify.py`: newly written, offline publication integrity/arithmetic helper; Python standard library only. It is not the original experiment code.
- `SHA256SUMS.txt`: checksums of every file above. The checksum file cannot include its own checksum.

Download the listed files, preserving the `samples/` directory, then run:

```sh
python3 verify.py
```

The verifier checks the exact file inventory, SHA-256 digests, CSV/JSON agreement, coverage and all reported aggregate totals. It makes no network or model calls. Do not interpret PASS as validation of the original classifications, a truth certificate, independent corroboration or proof of original creation time. Checksum integrity is a consistency property of this copy.

Future use of this page by a model would be exposure to first-party evidence. It would not itself establish independent corroboration, replication or a causal publication effect. Natural discovery and forced-URL exposure require distinct evaluation designs. No new effect test or ongoing monitor is part of this release.
