Task-specific distillation evaluation
A protocol for comparing a compact task-specific model with its reference model. The evaluation records accuracy, abstention behavior, review load, latency, and cost under the same test conditions.
The other method briefs
Context position sensitivity
A controlled comparison of identical instructions placed at different positions in long contexts. Any result must identify the model, context length, run count, scoring rubric, and observed failure modes.
Review steps in multi-stage agent tasks
A protocol for testing whether structured review improves safe task completion. It separates raw completion from verified completion and records latency, cost, human corrections, and unresolved risk.
Retrieval quality and latency as context grows
A method for measuring retrieval relevance and response time as indexed context expands. Results must include the storage backend, index configuration, hardware, cache state, and query mix.
Prompt injection resistance for tool-using agents
A proposed attack suite for agents that read external content or call tools. The protocol distinguishes detection, containment, incorrect tool use, data exposure, and cases requiring human review.
Synthetic training sample quality review
A planned rubric for assessing generated training samples before they enter a fine-tuning set. It evaluates provenance, duplication, task coverage, label confidence, and performance on a held-out real-world set.
The gates a result must pass
These are the minimum gates a future result would need to pass before it could support a recommendation.
A clear hypothesis
Each brief starts with a testable question tied to a real operating decision.
Documented inputs
Data sources, preparation steps, sampling methods, and exclusions must be recorded.
Repeatable methods
The model, configuration, scoring rubric, and run conditions must be specific enough to repeat.
Qualified reporting
Results must include limitations, failure modes, and the conditions where the conclusion does not apply.
