Meituan
OngoingMulti-turn agent RL · Livestream Q&A
Responsible for post-training a digital human livestream commerce Q&A agent powered by an in-house MoE model. Built multi-turn agent RL environments with verl, using GRPO and related methods including GiGPO to train ReAct-based reasoning and tool use.
- Designed reward components for output format, response length, trajectory turn count, and retrieval of relevant RAG information.
- Implemented separate product and store retrieval tools, combining keyword/BM25 retrieval with vector search to supply relevant information for answers.
- Connected multi-turn information gathering to livestream Q&A, with average end-to-end response time under 2 seconds.
Alibaba
2025.06 – 2025.09Qwen assistant · Post-training & planner rewards
Contributed to Qwen assistant development and Qwen3 post-training for task planning and tool use. Used verl and GRPO for reinforcement learning, with responsibility for reward design for the Planner module.
- Contributed to multi-agent workflows spanning intent clarification, complex task decomposition, planning, and tool use across services.
- Worked on synthetic query generation, data filtering and distillation, and multi-source RL rewards using LLM-as-judge.
- Supported YaRN-based context extension and decoupled training/inference deployment, including routing scripts adapted to Alibaba’s infrastructure.
- Enabled 4B and 32B models to replace Qwen-Max in several modules.
360
2025.03 – 2025.05Visual reasoning · Post-training & evaluation
Built 400,000 image-text training examples spanning math, code, and open-ended QA. Applied SFT, DPO, and GRPO, designed reference-based RL rewards, and developed automated business-task evaluations with OpenCompass.
AMD
2023.01 – 2023.06Algorithm intern · Assisted driving
Developed vehicle pose estimation and parking-space detection and tracking for assisted parking, with deployment and testing on a modified BYD Han EV.