Research archive · bounded studies, 29 September–7 October 2026
Research, with its limits visible
A dated record of what is documented, what we observed and what we have not tested. This is an evidence index—not a popularity ranking.
7 October 2026 · configured-marketplace Codex CLI documentary diagnostic
Installed/enabled, but the capability is missing: inspect the boundary
Preserve the failure, inspect source/state, then only the skill or current-session MCP branch. Presence and actual useful output remain separate.
- Practical output
- One short command block and four linear steps, supported-owning-control or redacted escalation; no new kit.
- Documentary scope
- Four undated primary pages inspected 12:56:43–12:57:06 UTC. Concrete inspection handles/field limits newly inspected; new-session prerequisite and separate quality/auth gates reconfirmed. Changed with evidence: none; no universal cure or cross-surface transfer.
- Not performed
- No CLI command, installation/toggle/restart, selection/tool invocation, actual output, diagnosis, repair or retest. Loaded provenance UNKNOWN; ten host actuals NOT RUN. Prior reports/assets, 33/49 engineering evidence and dates retained.
6 October 2026 · public-channel gate + active-advice correction
Can a hook-bearing local bundle enter public submission?
Not as-is under the current documented ZIP restriction. Local support and optionality do not establish public eligibility.
- Practical output
- Early gate and one channel-choice card; 5 October core-plus-helper advice/copy now conditionally local, historical report unchanged.
- Documentary scope
- Submission, Packaging, Errors and bounded conversion sections. Explicit hook/registered-mapping exclusion newly inspected; conversion submission-path hook adaptation remains in genuine tension. No exception, source precedence, changed rule or observed rejection established.
- Not performed
- No separate useful public variant, implementation, adapted ZIP, host operation, portal parsing, scan/review/approval or media test. 33/49 offline evidence and ten host NOT RUN records retained; only channel claims advance.
5 October 2026 · documented prerequisites + hypothetical dependency design
Can a public plugin rely on a local startup hook?
Reject an installation-to-compliance promise. Put required policy/input handling in the proposed core and keep a local SessionStart helper optional; narrow, clarify, redesign or stop when essential local/enforced behavior is required. 6 October correction: that hook-bearing design is conditional local scope, not an eligible public ZIP.
- Practical output
- One before/after, compact non-executable dependency sketch, truthful release copy and added questions in the existing architecture worksheet.
- Documentary scope
- Five official pages' hook/dependency sections. Selected settings and cloud-Work exclusion reconfirmed; discovery/trust/orchestration/coverage and conversion-guide caution newly inspected. Changed with evidence: none. Local/cloud/synced qualification is not a full surface refresh.
- Not performed
- Entire fictional review/resources/stop procedure hypothetical / NOT RUN. No actual policy, host, helper, installation, trust result, model output, enforced gate or binary approval. Future evidence UNKNOWN / NOT RUN; no new kit/matrix/widget.
4 October 2026 · source-backed contracts + original hypothetical design
Separate private-data summaries from external sends
Do not bury optional send inside a read-only summary. Split and enforce the actual human consent boundary, ship read/unsaved-preview only, or obtain precise clarification.
- Practical output
- Before → redesign, three non-executable contracts, exact fictional confirmation and an inspectable provider-owned human-only approval design.
- Current documentary scope
- Four official pages plus only the error reference's paired annotation entries. Reconfirmed principles; finer contract/persistence/retry precision newly inspected. Changed with evidence: none. Annotation-field conflict narrowly checked 4 October, unresolved.
- Not performed
- Entire workflow/protocol/copy hypothetical / NOT RUN. No host/provider/send, approval, scan/review, new tool/kit or safety result. Incidental logging scope, actual approval fidelity and provider recovery remain unresolved.
3 October 2026 · source-backed boundary + hypothetical design review
Can a plugin serve SaaS customers without selling upgrades?
One fictional read-only project-status flow replaces an external digital upsell with linkless entitlement information. Identity, resource access and feature entitlement remain separate.
- Practical output
- Before → boundary → replacement, six source-mapped copy/next-decision cases, conditional destination review and business falsification questions.
- Current documentary check
- Five official pages, bounded sections. Existing SaaS boundary reconfirmed; saved-method API/UI detail newly inspected. No changed rule or release date established. Physical-goods checkout scope remains conflicting / unresolved.
- Not performed
- All six cases are hypothetical design review / NOT RUN—not observed tests or a compliance score. No live destination, account, host, API, checkout, payment, portal or support test; no new kit/widget or demand study.
2 October 2026 · change evidence, not a rollout claim
I changed my plugin. What actually changed for adopters?
An original four-state model and runnable offline recorder: reproduce a same-version instruction edit, then identify independent loaded-copy, public-package and live-backend evidence.
- Practical output
- Recorder, worked JSON and blank next-evidence worksheet; synthetic revisions, not new sample releases.
- Observed engineering
- 48 subprocess commands plus one invariant: 49 recorder assertions passed. Exact download independently reproduced. Not plugin behavior, safety or update delivery.
- Current documentary check
- Five primary resources; consequential rules reconfirmed, no substantive rule change/date established. Endpoint contradiction persists; pending-review retention is not guaranteed by the inspected text.
- Not performed
- Host/session, public portal/listing and remote deployment/authorized behavior tests: all NOT RUN, observed fields UNKNOWN. Reader declarations never substitute for observations.
1 October 2026 · current documentation + exercised offline tooling
Offline release evidence is useful—but it is not a host result
An original runnable 19-file release-rehearsal kit for preserved Meeting Evidence 0.1.0: finite profile/ZIP checks, genuine failure → repair and ten separate blank host specifications.
- Observed engineering
- 32 subprocess commands plus one deterministic invariant: all 33 TOOLING assertions passed. Independent downloadable-kit reproduction succeeded. Not model performance or a safety result.
- Consequential clarification
- Portable schema, Codex logo/composerIcon validation and primary public icon are distinct scopes. The unchanged two-file sample is not established Codex-ready. No substantive policy change was established.
- Not performed
- Host acceptance, installation, activation/quality/security tests, portal validation/scans, review/approval/chosen plugin publication and authentic product media/rendering.
29 September 2026 · documentation + read-only observation
OpenAI plugins: package, access and publication
Current architecture/surface map, four task-selected examples, an original two-file package and its narrow static check, present policy constraints, and a qualified chronology.
- Established
- Package versus execution layers; current IDE exclusion; local/public milestone distinction; current digital-commerce restrictions and skills-only public update limit.
- Not established
- Universal-directory launch date; authenticated example behavior; exact account tiers/parity/scopes; review timing; search placement, adoption or revenue.
Evidence language
- Documented: a primary user/developer source describes the capability or rule.
- Listing observed: a read-only directory search returned an entry in this account context, not necessarily in yours.
- Offline engineering: observed checker/recorder commands, diagnostics and selected byte/structure differences—not model semantics, loaded/public bytes, deployment or safety.
- Reader declaration: intended channel, review/version/deploy label or service-change description supplied as context—not an observation.
- Host-tested: requires actual accepted package/install/invoke/output inspection with recorded conditions. None of our sample workflows qualifies; ten 1 October cases remain NOT RUN.
- Hypothetical design review / NOT RUN: proposed copy and flows grounded in documented constraints, not observed output, test success or platform approval.
- Analysis: our inference from documented control points, not a provider forecast or measured benefit.
What still needs a supported host
- One skills-only workflow and one authorized connected task; capture client/version/account, prompts, invocation, actual output and failures.
- Authentic redacted screenshots and a captioned explainer after useful text is available. No fabricated interface images.
- Recheck docs/catalog changes and obtain clarification on annotation fields and endpoint moves; investigate migration continuity only within the relevant account scope.
Links and prices may change. The checked date is not an announcement date or a checkout quote. Machine-readable retrieval guidance explains citation/availability conventions.
Historical route expansion · 29 September 2026
The 29 September 2026 study has been reorganized into ten developer-task routes. Its source/access/price check dates and saved observations are unchanged. That earlier expansion performed no new browsing, installation, invocation, migration or submission. The separate 1 October study rechecked nine consequential primary sources and added exercised offline tooling; it did not refresh catalog/pricing/migration/chronology evidence or run a product workflow.
- Documentation conflicts
- Annotation-justification and MCP endpoint-update guidance remain unresolved. Current IDE exclusion takes precedence over historical launch language; package variants and Canva entitlement limits stay explicit.
- Evidence to collect, not results
- Actual skill and connected-workflow records, precise client/account/version conditions, redacted outputs/screenshots and policy clarification remain future work.
Full architecture and four dated task records · Exercised offline checks and separate blank host contract · Submission conflicts and readiness
The 2 October study narrowly refreshed update/session/MCP claims and added original change-record tooling. In that 2 October study, annotation-justification retained its 1 October paired check; only the later 4 October annotation pairing advances it. Catalog, prices, migration, workspace sync and chronology retain their saved September dates. Current source comparison and exclusions.
The 3 October study narrowly inspected commerce and identity/resource/entitlement sections. General checkout wording and physical-goods saved-method API/UI guidance remain conflicting / unresolved, with no digital exception or universal enabling inferred. Payout/placement and other September observations were not refreshed; endpoint conflict stays 2 October; the later annotation-only pairing advances that conflict to 4 October without refreshing readiness. Current comparison and clock-bracket basis.
The 4 October study inspected action contracts, pure versus stateful effects, independent hints and server-enforced consent. Its 12:59:24–12:59:55 UTC clock bracket is not per-request HTTP time. Logging's incidental-infrastructure annotation scope and real human approval/status recovery remain unresolved; no new rule or implemented operation was established. Developer-first remains owner preference; disconnected analytics is not zero readership.
The 5 October hook/dependency check brackets current calls at 13:02:45–13:03:15 UTC, not HTTP timestamps or a rule-change date. Only H01–H09 advance: annotation conflict remains 4 October, endpoint 2 October, checkout 3 October; unrefreshed catalog/prices/migration/chronology/surface/worksheet claims keep saved dates. Current definition trust is not established all-byte dependency provenance; policy instructions are not reliable enforcement. Exact source headings and comparison basis.