GOVDOSS
Chapters
Main site
← All research

Evaluation note / 02

What a 20-case skill-routing review revealed

An internal routing review found a handoff error between architecture design and execution planning. The correction clarified who owns each stage.

In this publication

Key takeaway

Lifecycle state matters as much as domain. Once an architecture is approved, a request for an execution baseline should route to implementation planning.

What was reviewed

The internal July 2026 report describes 20 fixed prompts across strategy, research, security, programs, operations, learning, resilience, innovation, and governance. Each prompt was assigned an expected primary skill based on its dominant outcome and stage of work.

The review used installed skill names and descriptions to select one primary skill. It assessed routing and action boundaries. It did not evaluate the full analytical output of every selected skill or run a production workload.

Figure 01Reported evaluation method

How the routing review was structured

How the routing review was structured. Twenty fixed prompts supplied the cases. The comparison tested primary-skill selection, without evaluating every selected skill’s full output.

Twenty fixed prompts supplied the cases. The comparison tested primary-skill selection, without evaluating every selected skill’s full output.

Source: GovDOSS Skill Routing Benchmark — 20 Mixed Cases · internal report · July 28, 2026.

Read figure description

For each of 20 fixed prompts, the expected primary skill was assigned from the dominant requested outcome and lifecycle stage. The selected route came from installed skill names and descriptions. Comparing the expected and selected primary skill produced a match or mismatch for each prompt. This was a routing review, not a full-output evaluation.

Reported result

The report records 19 correct initial routes out of 20, or 95%. The one mismatch concerned an approved regulated-AI architecture whose next requested output was a 12-month execution baseline.

After the two skill descriptions were clarified, the report lists 20 of 20 cases as corrected. It explicitly documents a targeted rerun of the failed case and two adjacent controls. This is a correction result on known cases, not a new holdout test or an estimate of production accuracy.

The report records no external or consequential actions during the review. That finding describes this bounded exercise and does not establish runtime safety.

Figure 02Reported result · 20 prompts

Initial routing results on the fixed set

Initial routing results on the fixed set. Each cell is one prompt in the July 28 internal report. Cases 01–19 matched; case 20 did not. The corrected ledger and targeted rerun describe known cases, not an independent 20-case retest.

Each cell is one prompt in the July 28 internal report. Cases 01–19 matched; case 20 did not. The corrected ledger and targeted rerun describe known cases, not an independent 20-case retest.

Source: GovDOSS Skill Routing Benchmark — 20 Mixed Cases · internal report · July 28, 2026.

Read figure description

The initial report records 19 matches out of 20 routes, or 95 percent of this fixed set. Cases 01 through 19 matched, shown with circles. Case 20 mismatched, shown with a cross. After descriptions changed, the report lists a corrected ledger of 20 out of 20, while explicitly documenting a targeted rerun of three known cases: the failed case and two adjacent controls, all passing. There was no new holdout test. These counts do not estimate production accuracy.

The failure and the correction

The request already had an approved architecture and control boundary. It asked for workstreams, owners, milestones, resources, dependencies, acceptance criteria, and transition to operations. The initial route selected the regulated-AI design skill instead of the program-execution skill.

Both skill descriptions used overlapping planning language. The correction assigned unresolved hosting, boundary, and readiness decisions to the design skill, and execution baselines following approved decisions to the program-execution skill.

  • Boundary unresolved: route to architecture and deployment design.
  • Target state approved: route to execution planning when workstreams, owners, milestones, and transition are the requested deliverables.
  • Stage unclear: establish the decision state before choosing a primary skill.
Figure 03Routing rule from the correction

Route by the stage of work

Route by the stage of work. The requested deliverable matters alongside approval state. An approved architecture routes to program execution when an execution baseline is the requested outcome.

The requested deliverable matters alongside approval state. An approved architecture routes to program execution when an execution baseline is the requested outcome.

Source: GovDOSS Skill Routing Benchmark — 20 Mixed Cases · internal report · July 28, 2026.

Read figure description

Check the architecture and control-boundary decision state. If unresolved, route to architecture and deployment design to resolve hosting, boundary and readiness. If approved and an execution baseline is needed, route to program execution planning for workstreams, owners, milestones and transition. If the lifecycle stage is unclear, establish the decision state before selecting the primary skill. These are alternative branches, not sequential steps.

What the result does not establish

Correct routing does not prove that the chosen skill produces a complete or reliable answer. The report separately identified template content in one skill body, despite a correct metadata-routing result.

The publication is a summary of the internal report. Raw execution traces, independent scoring, repeat trials, and a held-out evaluation set are not included. The result should be read as a diagnostic finding about descriptions and handoffs.

Figure 04Evidence scope matrix

What the available evidence supports

What the available evidence supports. A reported route match is a narrower finding than answer quality, runtime safety or performance on unseen requests.

A reported route match is a narrower finding than answer quality, runtime safety or performance on unseen requests.

Source: GovDOSS Skill Routing Benchmark — 20 Mixed Cases · internal report · July 28, 2026.

Read figure description

Primary-skill routing was reported on 20 fixed cases. Complete output quality is not established by correct routing; the report found template content in one skill. Runtime safety is not established by this exercise. Generalization is not established because no held-out evaluation is included. These are evidence boundaries, not quality scores.

Limitations & next step

Internal, fixed-case routing review; no independent validation. Corrected cases were already known. The findings do not establish end-to-end quality, security, generalization, or production readiness.

Next research step

Preserve the original cases for regression checks and add unseen short, compound, and lifecycle-ambiguous prompts. Evaluate actual outputs separately from routing.

Figure 05Planned evaluation design

Separate regression from unseen-case evaluation

Separate regression from unseen-case evaluation. Keep the original 20 cases for regression checks and add unseen prompts. The new set’s size and acceptance thresholds remain to be defined.

Keep the original 20 cases for regression checks and add unseen prompts. The new set’s size and acceptance thresholds remain to be defined.

Source: GovDOSS Skill Routing Benchmark — 20 Mixed Cases · internal report · July 28, 2026.

Read figure description

Maintain a regression set of the original 20 known prompts. Add a separate unseen evaluation set with short, compound and lifecycle-ambiguous prompts; its size is not specified. For both sets, score routing and actual answer quality separately. Report known-case and unseen-case results separately. This is a planned study with no new results.

Questions about this research? Contact GovDOSS.