Project Dashboard

Operational Agency Rules → LLM Labeling → Scenario Benchmark

A single-page view of the full pipeline: rule development, interpretive-guide refinement, data cleaning, human labeling, LLM-assisted scaling, behavioral target generation, Bloom-based scenario construction, and runtime evaluation.

Current overall progressApprox. 72%
422Clean adjusted cases
65Development cases
25Held-out validation set
~332Remaining cases

Where the project stands now

Phase 1

Operational ruleset development

Three human cross-labeling rounds, disagreement analysis, and consolidation into the final active rules.
Completed
Phase 2

Interpretive guide refinement

Boundary conditions, displacement logic, and merged-rule guidance synchronized with the final taxonomy.
Completed
Phase 3

Adjusted fact cleaning

Conclusion leakage removed; original source fact and conclusion retained privately for provenance.
Completed
Phase 4

LLM labeling setup

Dev, held-out, few-shot, and remaining-case files organized; experiment harness being built.
In progress
Phase 5

Behavioral target generation

Trigger condition, compliant behavior, and violation behavior after final rule labels are locked.
Upcoming
Phase 6

Bloom scenarios + benchmark

Scenario kernels, modifiers, expert review, rubrics, and runtime evaluation.
Upcoming

Snapshot

Human coding rounds
3
40 → 25 → 25 cases
Held-out reliability
Jaccard ≈ .80
Post-freeze validation
Current focus
LLM harness
Zero-shot → 4-shot → 8-shot
Private provenance
Preserved
Source fact + conclusion

End-to-end flow

1. Source layerRestatement illustration
Original fact + original conclusion
2. Adjusted layerAI-adapted factual scenario
Conclusion leakage cleaned
3. Human rule labelingCross-labeling + adjudication
Final rules + guide
4. LLM scale-upZero-shot / few-shot dev eval
Human review of disagreement
5. Behavioral targetsTrigger condition
Compliant / violating behavior
6. Scenario kernelsAbstract factual structure
Rule-grounded benchmark seeds
7. Bloom scenariosModifiers + expert review
Public benchmark instances
8. Runtime evaluationTask success vs duty compliance
Agent rollouts + grading

Completed milestones

  • Final operational taxonomy frozen.
  • Interpretive guide synchronized with merged rules.
  • Adjusted facts cleaned for conclusion-blind coding.
  • 65-case development pool separated from 25-case held-out set.
  • Few-shot candidates selected to cover major rule boundaries.
  • Original Restatement material retained privately rather than publicly released.

Current to-do stack

  • Finish provider-agnostic LLM labeling harness.
  • Run a zero-shot pass on the dev set first.
  • Compare zero-shot, 4-shot, and 8-shot prompts.
  • Freeze prompt and model configuration.
  • Run the 25 held-out cases once for final evaluation.
  • Scale-label remaining cases with targeted human verification.

Phase details

PhaseGoalOutputStatus
Rule developmentTranslate doctrine into operational duties.Final active rulesetDone
Interpretive guideResolve rule boundaries.Final coding guideDone
Data cleaningPrepare conclusion-blind factual inputs.Clean adjusted factsDone
Human validationAssess reliability and adjudicate labels.65 dev + 25 held-outDone
LLM setupDevelop reproducible annotation pipeline.Zero/few-shot harnessNow
Scale labelingLabel remaining cases with LLM + humans.Full rule-labeled corpusNext
Behavioral targetsConvert labels into observable conduct.Trigger/compliant/violation fieldsNext
Benchmark releaseGenerate Bloom-derived scenarios and rubrics.Public benchmarkLater
Recommended next action
Finish the annotation harness and run a zero-shot dev pass before adding few-shot examples. Only after the prompt/model setup is frozen should the held-out 25 cases be evaluated.