Context
In April 2026, the protocol author used the deep-research features of Google (Gemini Deep Research), OpenAI (ChatGPT Deep Research), and Anthropic (Claude Opus 4.6 Thinking) to generate analyses of the Agentic Reasoning Protocol. The outputs were produced on the author's request; they are not independent reviews and not peer review.
The outputs named open research questions that need independent investigation. AI-generated agreement with the protocol is not used here as evidence for it.
This page consolidates those questions into a formal, open research agenda. Contributions are welcome via GitHub Issues.
RQ1: Standardized Evaluation Benchmarks
Source: ChatGPT Deep Research
Do AI-generated responses improve measurably when a domain's
reasoning.json is present in the retrieval context?
Proposed Methodology
- Select N domains across verticals (SaaS, consulting, e-commerce, healthcare)
- Generate baseline AI responses about each entity without ARP
- Deploy
reasoning.jsonwith sourced corrections and context - Re-query after indexing and measure: hallucination rate, factual accuracy, entity attribution correctness
- Automated fact-checking against
evidence_urlreferences
Open Sub-Questions
- Which AI platforms (Perplexity, ChatGPT, Gemini, Claude) show the strongest ARP responsiveness?
- Does the Pink Elephant Fix demonstrably outperform traditional negation-based corrections?
- What is the minimum indexing latency before ARP corrections take effect?
RQ2: Independent Experiment Replication
Source: ChatGPT Deep Research
Can the Ghost Site experiment, Canary Token forensics, and Citation Tracking results be independently replicated by third parties?
Experiments to Replicate
The findings below are reported by the protocol author. The raw data are not published in the ARP repository, and the findings have not been independently replicated.
| Experiment | Finding reported by the author | Replication Needs |
|---|---|---|
| Ghost Site | Dominant AI source within 24h | New domain, structured data only, multi-platform query |
| Canary Tokens | GPT/Gemini ingest reasoning.json | Unique tokens per platform, automated monitoring |
| Citation Tracking | 0% → 67% across 6 platforms in 22 days | Standardized query set, daily measurement |
| Zero Hallucination | Controlled ChatGPT case study | Multiple LLMs, statistical significance |
Counter-observation, also reported by the author: on 2026-05-13 the Vercel edge logs of
arp-protocol.org showed no bot requests for /.well-known/reasoning.json; the file
was added to the sitemap afterwards
(commit 826b94d).
RQ3: IETF Standardization Pathway
Source: ChatGPT Deep Research, Gemini Deep Research
Which standardization pathway, if any, fits a .well-known URI
that serves an entity's self-description?
Current Status
- Individual Internet-Drafts on the IETF Datatracker, with no IETF stream or working group assigned:
draft-deforth-arp-00(DNS-bound signatures for v1.x; posted 2026-04-18, expires 2026-10-20) anddraft-deforth-arp-reasoning-protocol-00(ARP v2.0; posted 2026-04-28, expires 2026-10-30); submission does not imply IETF endorsement or working-group adoption - W3C AIVS Community Group introduction in progress
Open Questions
- Is an IETF RFC, a W3C Community Group Report, or both the better path for ARP?
- How should the protocol handle versioning across RFC iterations?
- What is the relationship between ARP and the emerging AI Verifiable Standards (AIVS)?
RQ4: Multimodal Extension
Source: ChatGPT Deep Research
Can the ARP schema be extended to describe non-text entities — images, video, IoT devices, autonomous vehicles?
Considerations
- Image agents: Can
reasoning.jsoncarry sourced corrections relevant to visual AI (e.g., product image misidentification)? - IoT agents: Can sensor-equipped autonomous systems read domain-hosted self-descriptions, for example device specifications?
- Video: Can self-description statements be temporally scoped (valid for specific content windows)?
RQ5: Trust Model Adversarial Analysis
Source: ChatGPT Deep Research, Gemini Deep Research
What are the attack surfaces of a self-attested reasoning file, and which of them does cryptographic signing (since v1.2) address?
Threat Vectors
| Threat | ARP v1.1 Mitigation | ARP v1.2/v1.3 Mitigation |
|---|---|---|
| False self-attestation | Good faith (same as schema.org) | Ed25519 signature attributes the file to the domain operator; it does not show that the statements are true |
| Man-in-the-middle | HTTPS transport security | HTTPS + signature verification |
| Domain spoofing | DNS resolution | DNS TXT record binding |
| Competitor sabotage | Ethics policy | Signature attribution + community reporting |
| Altered copies (e.g., extended expiry date) | — | Signature covers the metadata (enveloped pattern); payload-only legacy signatures are rejected (v1.3) |
| Compromised or retired key | — | New selector per key; old selector revoked with an empty p= (v1.3) |
RQ6: Long-Term Search Impact
Source: ChatGPT Deep Research
What is the long-term impact of ARP on AI search results? Does the effect persist, amplify, or decay over time as AI models retrain?
Measurement Dimensions
- Citation persistence: Do AI platforms continue citing reasoning.json after model updates?
- Training integration: Do ARP self-descriptions eventually enter model training data?
- Competitive dynamics: When multiple entities in a vertical deploy ARP, how do AI systems resolve conflicting claims?
How to Contribute
This research agenda is open. AI researchers, RAG engineers, and domain owners are invited to contribute:
- Replicate experiments — Run the Ghost Site or Canary Token experiments independently and share results
- Propose benchmarks — Define standardized evaluation datasets via GitHub Issues
- Submit findings — Formal research contributions welcome via GitHub Issues or as independent publications
- Build integrations — LlamaIndex, CrewAI, AutoGen loaders welcome via Pull Request