LLM Post-Training, Agentic RL, Self-Evolving Agents and Sales Agents
Overview:
We research on efficient LLM post-training methods for agentic-RL and self-evolving agents,
including error-driven learning from failed trajectories, fine-grained credit assignment,
progressive curricula for long-horizon RL, and router-aware MoE fine-tuning;
also built SOP-guided RL for conversational sales agents and realistic user simulation
for interactive agents. Research has been translating into production systems for
financial marketing, customer service, and compliance.
On-Policy Self-Distillation and Self-Evolution for LLM Agents:
Turn failed trajectories into actionable supervision by clustering recurring failures
into structured mistake abstractions, locating the agent’s error frontier on long-horizon
tasks, and converting validated failures into frontier-aligned skills through on-policy
self-distillation. The combination of external error memory and recursive skill internalization
enables continual improvement without external teachers or additional rollouts.
Fine-Grained Credit Assignment for Long-Horizon Agentic RL: Extract learning signals
from disagreement among failed rollouts and from the geometry of internal reasoning dynamics.
Task outcomes provide the global objective, while cross-rollout contrasts and manifold-based
signals localize errors and redistribute credit across tokens and turns, improving stability,
interpretability, and sample efficiency.
Efficient LLM Post-Training: Improve MoE fine-tuning through router-aware adaptation
that better exploits task-relevant experts with substantially fewer trainable parameters,
and accelerate reasoning RL through ability-anchored progressive curricula with small,
repeatedly re-anchored difficulty increments.
SOP-Guided and Persuasive Dialogue Agents: Develop reinforcement-learning methods
for financial outbound agents using structured SOP process rewards and profile-conditioned
user world models, enabling reliable multi-stage workflow execution, intent probing,
objection handling, and adaptive persuasion.
Profile-Driven User Simulation: Build controllable user simulators that combine
demographic/personality profiles with business intents derived from real call logs.
Fine-tuned on voice dialogues and optimized with hierarchical GRPO rewards,
the simulator supports realistic evaluation, intent-recognition training,
and cold-start deployment.
Publications:
Sell More, Play Less: Benchmarking LLM Realistic Selling Skill.
Xuanbo Su, Wenhao Hu, Le Zhan, Yuting Xie, Kailin Lyu, Kaijie Chen, Ziwei Li, Haibo Su, Yunzhang Chen, Ling Huang.
In Proceedings of EMNLP 2026.
[pdf]
[Project Homepage]
Mistake Notebook Learning: Batch-Clustered Failures for Training-Free Agent Adaptation.
Xuanbo Su, Yingfang Zhang, Hao Luo, Xiaoteng Liu, Ling (Leo) Huang.
In Proceedings of ACL 2026 Findings.
[pdf]
[Project Homepage]
Profile-Driven User Simulation via Reinforcement Learning with Hierarchical Rewards for Task-Oriented Dialogues.
Juan Li, Fangshu Chang, Ling Huang.
Under Submission.