A single-page view of the full pipeline: rule development, interpretive-guide refinement, data cleaning, human labeling, LLM-assisted scaling, behavioral target generation, Bloom-based scenario construction, and runtime evaluation.
| Phase | Goal | Output | Status |
|---|---|---|---|
| Rule development | Translate doctrine into operational duties. | Final active ruleset | Done |
| Interpretive guide | Resolve rule boundaries. | Final coding guide | Done |
| Data cleaning | Prepare conclusion-blind factual inputs. | Clean adjusted facts | Done |
| Human validation | Assess reliability and adjudicate labels. | 65 dev + 25 held-out | Done |
| LLM setup | Develop reproducible annotation pipeline. | Zero/few-shot harness | Now |
| Scale labeling | Label remaining cases with LLM + humans. | Full rule-labeled corpus | Next |
| Behavioral targets | Convert labels into observable conduct. | Trigger/compliant/violation fields | Next |
| Benchmark release | Generate Bloom-derived scenarios and rubrics. | Public benchmark | Later |