Evaluation methods before conclusions

Neso AGI research noteReviewed September 2026Protocol draftAll notes

Draft methods for testing model performance, agent behavior, efficiency and safety. These briefs define what must be measured before the lab publishes a result.

Task-specific distillation evaluation

A protocol for comparing a compact task-specific model with its reference model. The evaluation records accuracy, abstention behavior, review load, latency, and cost under the same test conditions.

The other method briefs

Context position sensitivity

A controlled comparison of identical instructions placed at different positions in long contexts. Any result must identify the model, context length, run count, scoring rubric, and observed failure modes.

Review steps in multi-stage agent tasks

A protocol for testing whether structured review improves safe task completion. It separates raw completion from verified completion and records latency, cost, human corrections, and unresolved risk.

Retrieval quality and latency as context grows

A method for measuring retrieval relevance and response time as indexed context expands. Results must include the storage backend, index configuration, hardware, cache state, and query mix.

Prompt injection resistance for tool-using agents

A proposed attack suite for agents that read external content or call tools. The protocol distinguishes detection, containment, incorrect tool use, data exposure, and cases requiring human review.

Synthetic training sample quality review

A planned rubric for assessing generated training samples before they enter a fine-tuning set. It evaluates provenance, duplication, task coverage, label confidence, and performance on a held-out real-world set.

The gates a result must pass

These are the minimum gates a future result would need to pass before it could support a recommendation.

A clear hypothesis

Each brief starts with a testable question tied to a real operating decision.

Documented inputs

Data sources, preparation steps, sampling methods, and exclusions must be recorded.

Repeatable methods

The model, configuration, scoring rubric, and run conditions must be specific enough to repeat.

Qualified reporting

Results must include limitations, failure modes, and the conditions where the conclusion does not apply.

Questions about the protocol library.

Are these published research findings?

No. These are draft and planned evaluation protocols. A method brief should not be cited as evidence of a result. Completed findings will be labeled separately with their data, conditions, and limitations.

Can I reference a protocol?

You can link to a protocol as a description of a proposed method, but it should not be represented as a completed study or validated outcome.

When will results be published?

There is no fixed publication schedule. A result would only be labeled published if its protocol were complete, its evidence reviewable, and the limits of its conclusion stated clearly.

Do you collaborate with external researchers?

Use the contact page to propose a protocol review or dataset contribution. We will confirm whether a scoped evaluation is feasible.

What tools and frameworks do you use?

The stack depends on the protocol. Any completed result must identify the models, frameworks, evaluation harness, configuration, and relevant infrastructure used for that run.

Help review a protocol.

Bring a research question, dataset or evaluation concern. We can define the evidence needed before anyone treats an outcome as reliable.