Chenyu Wang

Chenyu Wang (王晨瑜)

Multimodal Agent Algorithm Intern · Shanghai AI Laboratory

Third-year M.S. student · SIMIT, UCAS

Research Interests: LLM Post-training, AI Agents, Multimodal Understanding

About Me

I am a third-year M.S. student in Electronic Information at the Shanghai Institute of Microsystem and Information Technology (SIMIT), University of Chinese Academy of Sciences (UCAS), advised by Baoqing Li and Min Tu. I am currently a Multimodal Agent Algorithm Intern at Shanghai AI Laboratory, working on LLM post-training, multi-agent collaboration, and retrieval-augmented decision-making.

My work spans rule-supervised dialogue synthesis, full-parameter supervised fine-tuning, exploratory reinforcement learning, and interactive evaluation. I also work on long-video understanding and description generation, and on security review for medical Agent plugins. My earlier research at XLab focused on domain adaptation for gas-sensor drift compensation.

Internship & Projects

Shanghai AI Laboratory

2026.03 — Present

Multimodal Agent Algorithm Intern · Multimodal Frontiers

1. LLM Post-training: Guideline-supervised Triage Policy Learning

2026.03 — Present
  • Constructed a rule-supervised multi-turn dialogue synthesis pipeline using 70 guideline rules and 3,000 training dialogues. Performed full-parameter SFT of a 9B model with class resampling and nurse-token supervision, and explored Outcome GRPO policy optimization.
  • On 201 internal evaluation cases, SFT increased operational-reference agreement from 61.7% to 74.1% and emergent-case recall from 9.5% to 69.0%. A second training seed and a second patient simulator reproduced the improvement trend.

These results measure agreement with an internal operational reference, not clinical validation.

2. MDT Agents: Planning and Retrieval-augmented Collaborative Decision-making

2026.06 — Present
  • Designed a nine-specialty multi-agent architecture for colorectal cancer multidisciplinary team (MDT) decision-making. Dual-source RAG retrieves patient history and clinical guidelines to support independent specialty reasoning, question planning, targeted retrieval, and final synthesis.
  • Organized 3,275 consultations from 1,865 patients at Zhongshan Hospital into longitudinal histories and a patient-disjoint evaluation dataset. Incorporated XGBoost / TreeSHAP statistical references and used paired ablations to evaluate decision consistency and error correction.

Ongoing research; blinded expert evaluation is pending.

3. DSH-med-plugin: Medical Agent Plugin Security & Infrastructure

2026.08 — Present
  • Conducted static security review of Plugin / Skill / MCP / CLI components for medical Agents. Organized provenance, versions, licenses, and review records for approximately 190 candidate repositories to support tool admission.
  • Completed static evaluation covering 11,126 entries and 12,830 components, plus sandbox validation of 6,795 scheduling units. LLM-as-a-Judge-assisted security review, combining source code, static alerts, and execution evidence, is in progress.

4. Domain Adaptation for Gas-Sensor Drift Compensation

2024.09 — 2025.09

XLab research · SIMIT, UCAS

  • Proposed a single-stage end-to-end adaptation framework combining multi-level, multi-kernel MMD, logit fusion, and dual-level Center Loss for distribution alignment and class discrimination; studied cross-dataset transfer.
  • Achieved 81.05% average accuracy on a public gas-sensor drift dataset. This work led to two first-author manuscripts, UnifiedGas and CDDA-Gas.

Multimodal Competitions

ECCV 2026 SLoMO Workshop · Long-video Understanding & Description Generation

Team ucas_xlab · Top three on all four Public leaderboards, including two first-place rankings.

EvalAI Public Test Split snapshot: September 8, 2026. These are Public rankings, not final award results.

Public leaderboard results and methods
TrackPublic rankMetricMethod
MovieQA Main ↗#3Accuracy 63.75%Multi-scale temporal sampling, subtitle / ASR grounding, and question-type-aware multi-model consensus.
MovieQA Special ↗#1Accuracy 49.63%Frozen Qwen3-VL-8B with 8 / 64 / 128-frame temporal views and weighted Jaccard consensus selection.
AD Main ↗#2AD Score 54.03Anchor-protected candidate fusion, risk-based interval selection, and double-blind dual-judge agreement grounded in visual evidence.
AD Special ↗#1AD Score 54.04Zero-shot Qwen3.5-35B-A3B sparse MoE inference with temporal context, dense frame sampling, and duration-constrained generation.

Publications

[1]
[2]
UnifiedGas: End-to-End Unsupervised Domain Adaptation via Hierarchical Multi-Level Alignment for Drift-Robust Gas Classification, IEEE Sensors Journal, under review (First author)
[3]
CDDA-Gas: Cross-Dataset Domain Adaptation for Metal-Oxide Gas-Sensor Drift Compensation, ACS Sensors, submission in progress (First author)

Education

University of Chinese Academy of Sciences (UCAS)

M.S. in Electronic Information

Shanghai Institute of Microsystem and Information Technology (SIMIT)

Advisors: Baoqing Li and Min Tu

2024.09 — 2027.06 (expected)

Xi'an University of Technology

B.Eng. in Electrical Engineering and Automation

School of Electrical Engineering

2019.09 — 2023.07

Skills & Honors

Technical Skills

  • LLM training: Python / PyTorch, full-parameter SFT, synthetic training data, and exploratory GRPO.
  • Agents: RAG, multi-agent collaboration, automated evaluation, and plugin security review.
  • Tools & research: Proficient with Claude Code and Codex; paper reproduction and English scientific writing. CET-6: 504.

Selected Honors

  • Merit Student, University of Chinese Academy of Sciences, 2026
  • First-class Scholarship, Xi'an University of Technology, 2020 / 2021 / 2022
  • Third Prize for Innovation Achievements, Xi'an University of Technology, 2022
  • Outstanding Student Leader, Xi'an University of Technology, 2020 / 2021 / 2022