Wenlong Deng
Open to Collaboration and Internship
My name is Deng Wenlong (邓文龙), I recently defended my Ph.D. in the Electrical and Computer Engineering department at the University of British Columbia, co-supervised by Prof. Xiaoxiao Li and Prof. Christos Thrampoulidis. My research focuses on long-horizon reinforcement learning and reasoning for large language models, particularly how to improve agents’ ability to solve complex multi-step tasks through better trajectory generation, reward design, credit assignment, and test-time search. I am especially interested in multi-turn reasoning, tool-using agents, and long-horizon data generation, with applications including coding, search, and medical diagnosis. I have also worked at Meta, Amazon, Google, and TikTok, where I applied these research ideas to practical, large-scale systems.
Previously: I obtained my master’s degree in Electrical Engineering at EPFL in 2019, where I was fortunated been supervised by Prof. Alexandre Alahi. I received my bachelor’s degree in Electronic and Information Engineering (Honors) at UESTC in 2017.
Personal Highlights
-
Data-Centric AI: Scalable Data Valuation & Synthetic Data. Study how training data shapes model learning and develop scalable methods for data valuation, long-horizon reasoning data curation, and synthetic data. Representative works include Evolving Dynamics (arXiv 2026), For-Value (ACL 2026), GMValuator (ICLR 2025), CCD (arXiv 2024), and MedReason (ICML 2026 GFM Oral).
-
Long-Horizon RL: Off-Policy Data, Reward Modeling & Policy Optimization. Study the learning dynamics of long-horizon reinforcement learning, with a focus on off-policy data, exploration–exploitation, and reward modeling. Representative works include NTHR (NeurIPS 2025), LLDS (ICML 2026), THR (ICLR 2026), Pass@K Surrogate (TMLR 2026), and Directional Reward Alignment (ICML 2026 AIWILD).
-
Industry Deployment: End-to-End Long-Horizon Agent Systems. Built end-to-end agent pipelines spanning long-horizon trajectory generation, milestone-based reward labeling, and stable multi-turn policy training. Applied these systems at Amazon and Meta, translating research advances into scalable training infrastructure and production workflows.
News
| Sep 25, 2026 | Our paper Learning Dynamics of Continual Learning: A Unified View of Data Attribution, Forgetting, and Plasticity Loss is online! A continual learner repeatedly faces three questions: what to learn from, what an update may disrupt, and whether future learning remains effective. We connect all three through the same token-level interaction. |
|---|---|
| May 04, 2026 | Two papers are accepted by ICML 2026: one on training collapse in multi-turn reinforcement learning, and the other on mitigating attention distraction in vision-language models. Many thanks to all my collaborators for their support and contributions! |
| Apr 08, 2026 | Out For-Value is accepted by ACL Main 2026, where we delve into the learning dynamics of SFT and introduce a forward-only data valuation framework that enables scalable and efficient value estimation for both LLMs and VLMs. Code avaliable at github. |
Selected Publications
- Agent ReasoningLearning Dynamics of Continual Learning: A Unified View of Data Attribution, Forgetting, and Plasticity LossarXiv, 2026
- Agent ReasoningOn Group Relative Policy Optimization Collapse in Agent Search: The Lazy Likelihood-DisplacementICML, 2026
- Data ValueFor-Value: Efficient Forward-Only Data Valuation for finetuning LLMs and VLMsACL, 2026
- ReasoningToken Hidden Reward: Steering Exploration-Exploitation in GRPO TrainingICLR 2026 (ICML AI4Math Best Paper), 2025
- ReasoningOn the Effect of Negative Gradient in Group Relative Deep Reinforcement OptimizationNeurIPS, 2025
- ReasoningMedReason: Eliciting Factual Medical Reasoning Steps in LLMs via Knowledge Graphs2025* Equal Contribution
- EfficiencyDARE the Extreme: Revisiting Delta-Parameter Pruning For Fine-Tuned ModelsInternational Conference on Learning Representations (spotlight 5%), 2025
- Data ValueGMValuator: Similarity-based Data Valuation for Generative ModelsInternational Conference on Learning Representations, 2025* Equal Contribution
-
Unlocking the Potential of Prompt-Tuning in Bridging Generalized and Personalized Federated LearningThe IEEE Conference on Computer Vision and Pattern Recognition, 2024 - MedicalLESS: Label-efficient Multi-scale Learning for Cytological Whole Slide Image ScreeningMedical Image Analysis , 2024
- MedicalOn Fairness of Medical Image Classification with Multiple Sensitive Attributes via Learning Orthogonal RepresentationsIn Information Processing in Medical Imaging (Accept rate 25%) , 2023
Talks
- Give a Talk at Northeastern University CS7150 on on-policy and off-policy Distillation, thanks for Jiaji’s Invitation!
Service
- 2024: PC member of FL@FM-IJCAI’24 and FL@FM-ICME’24
- 2023-now: Reviewer for NeurIPS, ICLR, ICML, TMLR, CVPR, ECCV, AISTATS, AAAI, and MICCAI
- 2026: Reviewer for ICML Position Paper
