Wenlong Deng

personal_img.jpg

Open to Collaboration and Internship

My name is Deng Wenlong (邓文龙), I recently defended my Ph.D. in the Electrical and Computer Engineering department at the University of British Columbia, co-supervised by Prof. Xiaoxiao Li and Prof. Christos Thrampoulidis. My research focuses on long-horizon reinforcement learning and reasoning for large language models, particularly how to improve agents’ ability to solve complex multi-step tasks through better trajectory generation, reward design, credit assignment, and test-time search. I am especially interested in multi-turn reasoning, tool-using agents, and long-horizon data generation, with applications including coding, search, and medical diagnosis. I have also worked at Meta, Amazon, Google, and TikTok, where I applied these research ideas to practical, large-scale systems.

Previously: I obtained my master’s degree in Electrical Engineering at EPFL in 2019, where I was fortunated been supervised by Prof. Alexandre Alahi. I received my bachelor’s degree in Electronic and Information Engineering (Honors) at UESTC in 2017.

Personal Highlights

  • Data-Centric AI: Scalable Data Valuation & Synthetic Data. Study how training data shapes model learning and develop scalable methods for data valuation, long-horizon reasoning data curation, and synthetic data. Representative works include Evolving Dynamics (arXiv 2026), For-Value (ACL 2026), GMValuator (ICLR 2025), CCD (arXiv 2024), and MedReason (ICML 2026 GFM Oral).

  • Long-Horizon RL: Off-Policy Data, Reward Modeling & Policy Optimization. Study the learning dynamics of long-horizon reinforcement learning, with a focus on off-policy data, exploration–exploitation, and reward modeling. Representative works include NTHR (NeurIPS 2025), LLDS (ICML 2026), THR (ICLR 2026), Pass@K Surrogate (TMLR 2026), and Directional Reward Alignment (ICML 2026 AIWILD).

  • Industry Deployment: End-to-End Long-Horizon Agent Systems. Built end-to-end agent pipelines spanning long-horizon trajectory generation, milestone-based reward labeling, and stable multi-turn policy training. Applied these systems at Amazon and Meta, translating research advances into scalable training infrastructure and production workflows.

News

Sep 25, 2026 Our paper Learning Dynamics of Continual Learning: A Unified View of Data Attribution, Forgetting, and Plasticity Loss is online! A continual learner repeatedly faces three questions: what to learn from, what an update may disrupt, and whether future learning remains effective. We connect all three through the same token-level interaction.
May 04, 2026 Two papers are accepted by ICML 2026: one on training collapse in multi-turn reinforcement learning, and the other on mitigating attention distraction in vision-language models. Many thanks to all my collaborators for their support and contributions!
Apr 08, 2026 Out For-Value is accepted by ACL Main 2026, where we delve into the learning dynamics of SFT and introduce a forward-only data valuation framework that enables scalable and efficient value estimation for both LLMs and VLMs. Code avaliable at github.

Selected Publications

  1. Agent Reasoning
    Learning Dynamics of Continual Learning: A Unified View of Data Attribution, Forgetting, and Plasticity Loss
    Yi Ren , Wenlong Deng, Guanzhe Hong , and 1 more author
    arXiv, 2026
  2. Agent Reasoning
    On Group Relative Policy Optimization Collapse in Agent Search: The Lazy Likelihood-Displacement
    Wenlong Deng, Yushu Li , Boying Gong , and 3 more authors
    ICML, 2026
  3. Data Value
    For-Value: Efficient Forward-Only Data Valuation for finetuning LLMs and VLMs
    Wenlong Deng, Qi Zeng , Jiaming Zhang , and 5 more authors
    ACL, 2026
  4. Reasoning
    Token Hidden Reward: Steering Exploration-Exploitation in GRPO Training
    Wenlong Deng, Yi Ren , Yushu Li , and 4 more authors
    ICLR 2026 (ICML AI4Math Best Paper), 2025
  5. Reasoning
    On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization
    Wenlong Deng, Yi Ren , Muchen Li , and 3 more authors
    NeurIPS, 2025
  6. Reasoning
    MedReason: Eliciting Factual Medical Reasoning Steps in LLMs via Knowledge Graphs
    Juncheng Wu* , Wenlong Deng*, Xingxuan Li , and 12 more authors
    2025
    * Equal Contribution
  7. Efficiency
    DARE the Extreme: Revisiting Delta-Parameter Pruning For Fine-Tuned Models
    Wenlong Deng, Yize Zhao , Vala Vakilian , and 3 more authors
    International Conference on Learning Representations (spotlight 5%), 2025
  8. Data Value
    GMValuator: Similarity-based Data Valuation for Generative Models
    Jiaxi Yang* , Wenlong Deng*, Benlin Liu , and 2 more authors
    International Conference on Learning Representations, 2025
    * Equal Contribution
  9. cvpr_unlock.jpg
    Unlocking the Potential of Prompt-Tuning in Bridging Generalized and Personalized Federated Learning
    Wenlong Deng, Christos Thrampoulidis , and Xiaoxiao Li
    The IEEE Conference on Computer Vision and Pattern Recognition, 2024
  10. Medical
    LESS: Label-efficient Multi-scale Learning for Cytological Whole Slide Image Screening
    Beidi Zhao , Wenlong Deng, Zi Han , and 5 more authors
    Medical Image Analysis , 2024
  11. Medical
    On Fairness of Medical Image Classification with Multiple Sensitive Attributes via Learning Orthogonal Representations
    Wenlong Deng, Yuan Zhong , Qi Dou , and 1 more author
    In Information Processing in Medical Imaging (Accept rate 25%) , 2023

Talks

  • Give a Talk at Northeastern University CS7150 on on-policy and off-policy Distillation, thanks for Jiaji’s Invitation!

Service

  • 2024: PC member of FL@FM-IJCAI’24 and FL@FM-ICME’24
  • 2023-now: Reviewer for NeurIPS, ICLR, ICML, TMLR, CVPR, ECCV, AISTATS, AAAI, and MICCAI
  • 2026: Reviewer for ICML Position Paper