1. LLM Post-training: Guideline-supervised Triage Policy Learning
2026.03 — Present- Constructed a rule-supervised multi-turn dialogue synthesis pipeline using 70 guideline rules and 3,000 training dialogues. Performed full-parameter SFT of a 9B model with class resampling and nurse-token supervision, and explored Outcome GRPO policy optimization.
- On 201 internal evaluation cases, SFT increased operational-reference agreement from 61.7% to 74.1% and emergent-case recall from 9.5% to 69.0%. A second training seed and a second patient simulator reproduced the improvement trend.
These results measure agreement with an internal operational reference, not clinical validation.