Composition Evaluation Research Matrix

This matrix records the decision and evidence status for every source named in

Composition evaluation research matrix

This matrix records the decision and evidence status for every source named in #2118. It is based on the local corpus under `research-papers`; no source is silently treated as stronger or broader than its recorded quality assessment. The inspected corpus revision is `a5acbb9ca1c7428a2aff9f4e577b1a48c3d2ed5e`.

REFEvidence statusDecisionHarness consequence
REF-020, Tree of ThoughtsPeer-reviewed NeurIPS 2023; local GRADE HIGHResolve: adopt as a deliberate-search comparison rationale, not a universal graph claimCompare multi-path search with single-pass; record evaluation error, pruning risk, task fit, and additional compute
REF-021, ReflexionPeer-reviewed NeurIPS 2023; local GRADE HIGHResolve: adopt generate-evaluate-refine as a baseline with evaluator safeguardsMeasure false-positive evaluation, premature stop, retry count, and last-accepted recovery
REF-024, LATSPeer-reviewed ICML 2024; local GRADE HIGHResolve: adopt as reasoning/acting/planning prior art, not proof that Flow graphs improve qualityInclude acting/tool tasks, search cost, value/evaluator identity, and recovery
REF-1275, Multiagent DebatePeer-reviewed ICML 2024; local GRADE A-Resolve: adopt parallel/debate evidence with explicit limitsCompare independent candidates, shared/independent models, evaluator disagreement, context/cost, and convergence-not-correctness
REF-1453, Evaluation TraparXiv v1 methodological preprint; local GRADE BResolve for methodology onlyState the claim first, test proxy routes, use contrastive failure cases, and keep the empirical claim gate closed on conformance data
REF-1454, Self-Improvement Can Self-RegressarXiv v1 empirical preprint; local GRADE BDefer direct transfer: training-time RLVR collapse is not an inference-composition resultRetain trajectory, early-stop, peak-result, and multi-metric regression hypotheses for later provider studies
REF-1527, ZEBRAWorkshop/preprint; local GRADE B+Resolve as an engineering candidate, not a default allocatorAblate budgets and phases; report controller overhead, transfer limits, requested/realized allocation, and task screening
REF-1528, Token BudgetsSingle-author arXiv preprint with executable artifacts; local GRADE B+Resolve for budget-boundary engineeringTrack non-bypassable ownership, delegated/retried spend, duplicate receipts, enforcement layer, and provider-accounting trust
REF-1537, BudgetThinkerarXiv/OpenReview preprint; local GRADE B+Defer direct adoption because it requires model training and inference-engine changesAdopt telemetry distinctions: adherence versus utilization, natural completion versus forced cutoff, and quality-per-token/time curves

Further-investigation queue

REFQueue stateEvidence needed to advance
REF-020Resolved for harness designProvider-era replication on the fixed task suite before performance claims
REF-021Resolved for baseline designIndependent evaluator and false-positive-stop measurements
REF-024Resolved for comparison designTool-task replication and cost/recovery comparison against plain DAG/loop controls
REF-1275Resolved for parallel-policy designNewer-model replication, blinded preference, and convergence calibration
REF-1453Resolved for claim-gate methodologyEmpirical validation of discriminative benchmark conditions remains open
REF-1454Deferred for direct composition inferenceEvidence that its training-collapse mechanism transfers to inference-time refinement
REF-1527Resolved as experimental ablationIndependent reproduction, controller-overhead accounting, and broader task families
REF-1528Resolved for hard-budget telemetryProvider billing reconciliation and non-Rust enforcement validation
REF-1537Deferred for direct runtime adoptionClosed-provider equivalent or a governed trained-model/inference-engine integration

Evidence boundary

Peer-reviewed results support the existence and task-contingent usefulness of search, reflection, and debate patterns. They do not establish that AIWG's specific Flow graph implementation improves quality. The preprints contribute testable risks and engineering hypotheses with stronger hedging. The shipped synthetic records resolve harness conformance only; provider performance, human preference, and general product positioning remain `NOT RUN` or blocked.