Research report · 4 October 2026
Separate private-data summaries from external sends
A “read-only summary” cannot hide an optional send. Expose retrieval, pure unsaved preview and external send separately. Permit send only when the backend can authenticate the actual approving human and enforce one fixed action with current rights and usable approval state.
Split and enforce—or omit send
Our recommendation: split and enforce the send boundary; otherwise ship read + unsaved preview only, or obtain precise host/provider clarification before adding the action. Official Define tools · Map use cases to tools / Define each contract supports splitting differing permissions, risk and confirmation. The operational design is our synthesis.
Complete advertised capabilities determine independent hints; a harmless default does not erase a send mode. Pure previews differ from saved artifacts, jobs and outbound messages. Accurate hints do not enforce access or consent. Guidelines · MCP requirements → Tools → Correct annotation / independent exposure, MCP review · tool-scanning metadata / annotation mismatch.
One inspectable worked boundary
Open the before → redesign review: authorized selected-note read → in-memory unsaved preview → separate authenticated-human approval → separately stateful external send. Three non-executable contracts show exact inputs, effects, destination and fail-closed behavior. One fictional recipient/project/source/subject/body confirmation makes the disclosure concrete; native expandable detail keeps the actual consent mechanism inspectable.
- Read and preview: bounded private retrieval/in-process computation each conditionally uses readOnlyHint=true, destructiveHint=false, openWorldHint=false. Preview saves/sends nothing. A sensitive read is still not a safety result.
- Send: irreversible open-ended external email conditionally uses readOnlyHint=false, destructiveHint=true, openWorldHint=true. Saved approvals/dispatch records are writes, not pure preview. No value authorizes an operation.
- Human consent: proposed provider-owned human-only issuer, separate from model tools, authenticates the human and freezes displayed subject/account/project/resource versions/recipient/exact subject/body. Current rights, expiry and one-attempt state must still match at execution.
- Uncertainty: unknown/mismatched authority or changed arguments → no dispatch. Reserved attempt with unclear outcome blocks another send. Verified provider-specific idempotency/status evidence may support reconciliation, not guaranteed exactly-once delivery.
Entire workflow, contract fields/limits, source bundles, copy, step-up, fixed record, 10-minute example, reservation and recovery: original hypothetical design / NOT RUN. No implemented provider, new tool/kit/widget, OpenAI approval-token protocol or certification. The three traps are lessons—not tests. Source text, a write scope, account connection, chat text alone and model confirmed:true are not trustworthy consent proof.
Security & Privacy · Principles / Data handling / Prompt injection and write actions / Authentication & authorization supports server validation, least privilege, minimization and irreversible-action human confirmation. It does not establish our proposed channel or approve this design.
Bounded source comparison, not a policy-change claim
Runner read the supplied public snapshot and historical claim records before independent current primary retrieval. Four core official pages were opened; only annotation entries of the error reference were added for the conditional conflict pairing. Observed tool-call clock bracket: 4 October 2026, 12:59:24–12:59:55 UTC, not per-request HTTP timestamps or release dates. Pages are undated; earlier saved records are claim summaries, not complete archived subsections.
- Define tools · A01–A03
- Map use cases / Define each contract / Plan safety annotations: split boundaries and explicit inputs/output/authorization/effects/failure. Contract-level precision newly inspected; metadata-versus-authority principle reconfirmed.
- Guidelines · A04–A10
- MCP requirements → Tools: independent exposure, complete descriptions, explicit/independent booleans, pure-versus-stateful effects, minimization and retry risks. Broader truthful-metadata/minimization principles reconfirmed; precise mode/persistence/retry claims newly inspected, not newly released.
- Security & Privacy · A11–A12
- Principles / Data handling / Prompt injection and write actions / Authentication & authorization: saved consent, minimization and server-enforcement obligations reconfirmed. OAuth/network/UI guidance not refreshed; no resistance or enforcement test.
- MCP review · A13/A15
- Metadata stored during tool scanning / Review FAQs → annotation mismatch: accurate imported values reconfirmed; all-mode/indirect-effect precision newly inspected. No scan or portal operation.
- Annotation-error pairing · A14
- Final directory submission / annotations_required / justification_required: annotation entries only. Reconfirmed disagreement with Guidelines, not a full submission-readiness refresh.
Changed with evidence: none. Reconfirmed means a saved claim remains supported; newly inspected means finer claim-level precision absent from saved summaries, not a new rule. A01–A15 are our editorial references, not official rule IDs. Metadata checked/access times record saved-ledger review, with that basis—not individual fetches.
Three unresolved dependencies
- Annotation justification · paired 4 October scope
- Guidelines says explanations are no longer required; the error reference still requires them and review uses justification language. No authenticated form, support answer or precedence resolves the field requirement. Confirm the actual path's fields; explanations cannot override correct behavior values. Keep other readiness at 1 October.
- Logging scope
- Review's state-change examples include writing logs; Security advises redacted logs. The inspected texts do not delineate incidental infrastructure classification. Deliberate application persistence is not read-only; neither universal “every log write” nor “logs never matter” follows.
- Human approval and recovery
- No actual host/provider human-only channel, argument fidelity or status/idempotency semantics observed. The design requires those prerequisites; a deployment must prove them. Unknown outcome blocks another dispatch, not authorizes a retry.
Retained evidence and work not performed
Runtime confirms the 3 October guide release published. This run's candidate is separate and remains subject to Runtime acceptance/delivery/publication. Original Meeting Evidence/downloads, prior kits/logs/JSON/reports and accepted proposals remain intact: 33 checker and 49 recorder assertions are separate engineering observations; ten host cases and six commercial cases remain NOT RUN.
Endpoint contradiction stays checked 2 October; checkout scope tension stays 3 October. Catalog, prices, migration, chronology, workspace sync, payout/placement and unrefreshed worksheet claims retain saved dates. Developer-first is owner preference, not measured demand; analytics is disconnected, not zero readership.
No account/provider/host/API/send, installation/deployment, protected portal/support, scan/review/approval, destination audit, executable tool/package, new engineering suite or media operation ran. Future authorized synthetic-data, human approval, fixed-argument and redacted outcome evidence is prerequisite work—not reported test success. Manager editorial review, the sole local final-shell preflight and Runtime's repeat/host gates are separate.