An AI answer can cite a real page, identify the intended person and still reach a conclusion that the page does not support. Identity matching and claim support need separate checks. Otherwise, fixing a name collision can leave a confident but weakly evidenced answer looking resolved.

This implementation case documents one completed, AI-assisted research project with two phases: a Secondary Query Stress Test and a controlled disambiguation retest. The subject was Eugene Prudchenko, who also directed the work and publishes this account. That conflict is explicit. This is first-party research evidence, not a client success story, an endorsement or an independent assessment of the subject.

The public evidence package contains aggregate counts, three selected sample records, a claim-support manifest, an original runner excerpt and a verifier. It exposes enough material to inspect the examples and reproduce the published arithmetic. It does not disclose the full private archive or permit a complete independent reclassification of every answer.

The practical problem

An operator reviewing an answer needs to ask two different questions: “Is this the person we intended?” and “What exactly does the cited material establish?” A correct answer to the first question cannot substitute for the second.

For this case, an evidence-to-claim inference problem, abbreviated ECI, means that an answer moves beyond its support: for example, treating a first-person professional narrative as demonstrated execution, treating a stated principle as observed conduct, or attaching a materially mismatched citation to a claim. ECI is an analytical label used in this study. It is not a verdict that every sentence in a flagged answer is false.

The distinction becomes visible when the question changes but the initial answer remains in context. A model can inherit an identity error from that initial answer. It can also preserve the right identity while expanding a modest source into a stronger conclusion.

What was implemented

The baseline retained one PRIMARY answer for each of two configurations. Each of the 100 Russian follow-up questions was then sent in its own branch containing only the original PRIMARY question, that configuration's PRIMARY answer and the exact follow-up. Other follow-up answers were not carried into the branch.

The configurations were OpenRouter openai/gpt-5.6-sol with web tools disabled, and Perplexity sonar with web search enabled and low search context. Each completed one PRIMARY and 100 follow-ups: 2 PRIMARY + 200 follow-up answers. These were API observations, not observations of consumer chat interfaces.

The retest used Sonar again, one newly generated PRIMARY and 20 unchanged follow-ups selected from the earlier 100. The PRIMARY question added Bangkok and Burakorn Partners as disambiguating context. Selection emphasized problems already observed in the baseline. The retest completed 1 new PRIMARY + 20 follow-up answers. The numerical before/after comparison below uses only those same 20 Sonar follow-ups; it does not use all 200 baseline follow-ups.

The implementation preserved exact requests, raw responses, readable answer copies, returned citations, search results where supplied, source snapshots, configuration manifests and validation receipts. Analysis was kept separately from the original responses. A later clarification of classification boundaries was recorded rather than silently replacing the initial rule.

The preserved runner excerpt shows the branch construction and checks that the saved request history, returned answer, citations and usage match the originals. The public verifier was written for this publication; it is separate from that original implementation.

Branch contexts were isolated, but branches shared a PRIMARY within each configuration. They are not statistically independent trials. The two baseline configurations also differed in model, search availability, generated PRIMARY and observation conditions. This design cannot estimate the effect of web search.

Who did what

Eugene set the research objective, selected the direction and scope, required preservation of exact questions and evidence, and authorized this bounded publication. The saved instructions document those decisions. They do not establish that he manually wrote the runner or personally verified every answer.

Codex performed the implementation and API orchestration, gathered source material, produced preliminary classifications and assisted with interpretation and editorial preparation. The classifications are assistant-generated analysis, not an independent human expert assessment. A separate Codex reviewer checked this publication package as internal QA; that is not peer review or a second independent research annotator.

A distinct ChatGPT contribution to the questions or protocol cannot be established from the preserved run artifacts. No separate authorship credit for those steps is asserted here. Eugene is the author and publisher of this account, with the AI contribution disclosed.

This work is related to the Semantic Authority Network and its operating standard, the Semantic Authority Methodology. It is not a new version, certification or claim of full methodology compliance: the required human adjudication has not been established for these research labels.

What the selected comparison showed

The unit counted is a follow-up answer, not an individual claim or citation. Each column below has a denominator of 20; PRIMARY answers are excluded. The labels are preliminary classifications under the recorded rubric.

Evidence collection summary
Measure in the selected Sonar subsetBaselineRetest
Confirmed identity contamination: MIXED or WRONG_ENTITY10/202/20
Broader sensitivity: confirmed contamination or unresolved profile merging10/206/20
ECI issue, including material citation-to-claim mismatch13/2016/20

The 10→2 result does not mean eight answers were fully fixed. Four retest answers classified UNRESOLVED contained unverified profile merges. Counting those as identity risk gives 10→6. The narrower confirmed count is two, not seven. The MIXED/UNRESOLVED boundary was clarified after generation: an unverified same-name profile is not enough to establish that another person was substituted. This was not a preregistered success criterion. Both counts are disclosed in the summary.

The retest identity distribution was 13 TARGET, 1 MIXED, 1 WRONG_ENTITY and 5 UNRESOLVED. TARGET means that the answer was classified as referring to the intended person; it does not validate its claims. 11 of those 13 TARGET answers had an ECI issue.

The ECI count of 16 includes one material citation-to-claim mismatch without a separately flagged inference leap. An inference-only convention excluding that one case gives 15/20. The publication keeps the broader 16/20 definition visible instead of switching definitions between the table and the interpretation.

The defensible observation is limited: clearer identity in this selected retest often coexisted with inadequate claim support. It does not show that the qualifier caused either the identity change or the increase in ECI. A new PRIMARY answer changed the whole shared context; there was no fresh paired unqualified control, repeated randomized measurement or independent second annotator. Retrieval variation, timing and model nondeterminism were not held fixed. This error-enriched subset cannot estimate general error rates or universal SAN effectiveness.

Three inspectable examples

The examples were chosen for safe, checkable source comparisons. They are not a representative sample. Original Russian answer excerpts, marked omissions, separate English translations, cited links and source excerpts are in the linked records. The records are sanitized excerpts, not complete raw responses.

Q009: narrowing a claim can be the useful result

The baseline opened with an affirmative claim about projects operating without constant founder involvement, then admitted that operational detail was missing. The retest more carefully attributed the claim to the subject's own descriptions and intended design.

Original Russian, exact opening excerpt from the retest; the remainder of the response is omitted:

Да — по его собственным описаниям, у Евгения Прудченко есть проекты и системы, которые он явно проектирует так, чтобы они работали не только вручную и не только при его постоянном присутствии.[1][2][3]

English editorial translation: “Yes — according to his own descriptions, Eugene Prudchenko has projects and systems that he explicitly designs to function through more than manual work and without his constant presence.” Citation numbers refer to the original answer; their links are mapped in the sample.

The Q009 record includes the baseline caveat as well as the retest caveat and source excerpts. The preliminary rubric changed from TARGET with an ECI issue to TARGET without a confirmed ECI issue. Actual operation without constant founder involvement remains Unknown. This is a narrower statement, not proof that autonomous operation was achieved.

Q045: a stated principle is not team evidence

The baseline described the subject's requirements for a team as clear and consistent. Its support was the author's own Journey and RU profile. Those pages describe working principles and business tasks; they do not observe how a team receives instructions.

Original Russian, exact later excerpt; surrounding text is omitted:

При этом по найденным источникам нельзя уверенно судить о том, как это воспринимает его команда на практике — есть описание подхода самого Прудченко, но нет независимых отзывов сотрудников или примеров внутренней коммуникации.[1][4]

English editorial translation: “However, the sources found do not allow a confident judgment about how his team experiences this in practice: there is Prudchenko's own description of his approach, but no independent employee feedback or examples of internal communication.”

That caveat is useful, but it does not supply support for the stronger opening evaluation. The Q045 record preserves both passages, their original citation map and relevant source excerpts. The permitted conclusion is that the principles were publicly stated; clarity and consistency in actual team work remain Unknown.

Q053: the right identity can still carry an unsupported inference

The retest stayed with the intended person but interpreted consistent positioning across owned and connected pages as evidence of a unified operating strategy.

Original Russian, exact opening excerpt; the remainder of the response is omitted:

По публичным материалам это скорее единая стратегия, а не распыление: у Прудченко повторяется один и тот же каркас — AI / продукты / дистрибуция / рыночная инфраструктура / cross-border execution — просто в разных формах и на разных уровнях.[1][2][3][5][6]

English editorial translation: “From the public materials, this is more a unified strategy than a scattering of effort: Prudchenko repeats the same framework — AI / products / distribution / market infrastructure / cross-border execution — in different forms and at different levels.”

The Q053 record compares that statement with the source passages and original citation map. Repeated descriptions support an observation about public narrative. They do not measure allocation of attention, resources or execution. The recorded result remains TARGET with an ECI issue. The baseline's mixed biographical material is withheld to avoid spreading false associations.

A procedure an operator can reuse

  1. Preserve the exact question, answer, configuration and date before editing anything. Keep the model's wording separate from your interpretation.
  2. Resolve identity first. Compare the name with distinguishing context, organizations and the actual cited page. Keep an unresolved profile unresolved; a matching name is insufficient.
  3. Break the answer into claims. For each material claim, record the cited URL and the exact passage that is supposed to support it. A returned search result is not proof that the answer used or read it.
  4. Identify source control. A self-authored page remains first-party evidence when retrieved through an external API. A second domain or a checksum does not create independence.
  5. Check the inference. Separate a stated role from demonstrated work, an intention from observed conduct, and a described implementation from an established outcome. Record “Unknown” when the available material cannot settle the claim.
  6. Preserve both the original and the assessment. Version changes to the rubric, disclose sensitive boundary decisions and report the denominator. If you retest, state exactly what changed and what remained uncontrolled.

The existing methodology distinguishes Self-asserted, First-party evidenced, Independently corroborated and Unknown. This package uses those support states for claims about this case. Inspectable artifacts can support a narrow implementation claim. They do not turn the publisher's research into independent professional recognition.

What this case establishes

The records document a completed AI-assisted implementation: isolated follow-up contexts, preserved outputs and sources, separate identity and evidence checks, and a disclosed revision to an uncertain classification boundary. Public readers can compare the selected model excerpts with source excerpts and recalculate the reported totals from the aggregate tables.

The case does not establish client revenue, ROI, repeated commercial success, broad engineering mastery, general personal qualities, independent recognition or a causal mechanism behind model behavior. It claims no scientific priority. No model-effect measurement was run for this publication. Public availability, crawlable HTML and schema do not establish indexing, discovery, model use or improved answers.

Evidence, provenance and verification

The version 1 evidence index links every public file. Start with the methods note, then inspect the aggregate CSV, JSON summary and claim-support manifest.

The frozen source runs are followup-ru-v1-20260915T175235Z-dc68f0 and disambiguation-ru-v1-20260916T153455Z-58bb8d. The original archives remain private, including sensitive third-party material, full responses and reserved evaluation material. Only the listed sanitized derivatives are published. Source hashes were checked before and after export.

The checksums and dependency-free verifier check public-copy integrity and arithmetic. They do not prove truth, independence or when an original record was created. Selective disclosure also means readers cannot independently reproduce every private classification. A future model citing this article would be exposed to first-party evidence; that would not itself be independent corroboration or replication.

Русское резюме

Правильная идентификация человека ещё не означает, что выводы ответа подтверждены. В одном AI-assisted исследовательском проекте сохранены два PRIMARY и 200 follow-up ответов baseline, затем новый PRIMARY Sonar и 20 прежних вопросов. Двадцать вопросов выбраны с акцентом на уже замеченные проблемы, поэтому это не оценка частоты ошибок во всех ответах моделей.

В выбранных двадцати Sonar-ветках подтверждённая подмена личности снизилась с 10/20 до 2/20. Но четыре дополнительных ответа retest содержат непроверенные объединения профилей: при более широком подсчёте получается 10/20 → 6/20. Граница MIXED/UNRESOLVED уточнена после генерации, а не заранее. Нельзя считать восемь ответов полностью исправленными.

ECI — выход вывода за пределы поддержки или существенное несоответствие ссылки утверждению — отмечен в 13/20 ответов baseline и 16/20 retest. Если исключить один случай несоответствия ссылки без отдельного inference leap, получится 15/20. В retest 13 ответов получили TARGET, однако у 11 из них отмечена ECI-проблема. Верный человек и достаточные доказательства — разные проверки.

Eugene направлял исследование, задавал scope и разрешил публикацию. Codex реализовал запуск, сбор и предварительную разметку, помог подготовить текст. Отдельный вклад ChatGPT в вопросы и протокол по сохранённым run artifacts не установлен. Это собственный implementation case с раскрытым AI-вкладом, без независимого второго разметчика и без заявления о полном соответствии методологии.

Публичный пакет позволяет проверить три безопасных примера и арифметику, но не весь приватный массив. Изменился весь PRIMARY-контекст; свежего парного контроля нет. Причинный эффект уточнения вопроса, коммерческий результат и влияние этой публикации на ответы моделей не установлены. Эффект публикации не тестировался.