What we're testing right now

Neso AGI research noteReviewed September 2026Protocol draftAll notes

Research questions, evaluation boundaries and protocols under review. Results are published only when the run data and scoring method are ready to inspect.

Protocols under review

Multi-agent task routing

Define how specialized agents should receive subtasks, expose their source context, and stop for approval. The protocol compares routing quality, latency, and operating cost.

Distillation efficiency

Specify a repeatable comparison between a compact task model and its larger reference model across classification, summarization, cost, and failure cases.

DomeAI retrieval latency

Map retrieval quality and response time as a private context store grows. Results will be reported only with the dataset, index settings, and hardware attached.

Long-context faithfulness

Test whether instruction placement changes compliance in long contexts, using the same tasks and scoring rubric across each model under review.

Synthetic data quality scoring

Design a scoring rubric for generated training examples before any sample is admitted to a task-specific training set.

Agent self-correction loops

Compare a baseline agent with a review-step agent while keeping tasks, tool access, and scoring constant. Human approval remains part of the test boundary.

How a completed experiment should run

A completed experiment should follow a structured process so its evidence can be reviewed and repeated.

  1. Hypothesis: start with a clear, testable question. What are we trying to prove or disprove?
  2. Design: define metrics, control variables, dataset size, and success criteria before running anything.
  3. Execute: run the experiment with logging at every step. Capture raw data, not just conclusions.
  4. Analyze and document: record positive and negative results under the same reporting standard once the evidence is reviewable.

Questions the protocols must answer

How small can a task model be before error cost outweighs inference savings?

The protocol records accuracy, abstention behavior, review load, latency, and cost together so a cheaper model is not treated as a win by default.

Which context placement produces the most faithful output?

Each placement uses the same instructions and tasks. Any reported difference must include the model, context length, run count, and scoring method.

When does an agent review step reduce errors enough to justify the latency?

The evaluation separates task completion from safe completion and records how often a human still needs to correct or reject the prepared action.

Questions about the experiments.

Can I access the raw data from your experiments?

This page currently publishes protocol summaries, not completed datasets or peer-reviewed findings. When a result is released, the model, task set, scoring method, and relevant limitations should be attached.

How do you decide what to experiment on?

Research questions are selected from workflow risks that can be measured: retrieval quality, model fit, review load, latency, cost, and safe handoff behavior.

Can I propose an experiment?

Use the contact page to propose a specific question about model behavior, agent systems, or training methods. We will confirm whether a scoped protocol review is feasible.

How long does a typical experiment run?

Timing depends on the protocol, available data, review requirements, and whether a result can be repeated. We do not assign a publication date until those conditions are clear.

Do you publish negative results?

A completed protocol should document negative and positive results under the same reporting standard. This page does not currently claim a published result set.

Have an experiment idea?

Bring a specific hypothesis, dataset or workflow risk. We can define the evidence and review gates a scoped evaluation would need.