SPaCe — Self-Paced Curriculum for LLMs
Self-paced curriculum learning that combines cluster-based data reduction with a multi-armed bandit sampler, matching baselines while using up to 100× fewer training samples.
ACL 2026
My name is Van Dai Do. I build efficient & safe LLMs—reinforcement learning, activation steering with memory, and retrieval-guided decoding. Additionally, I work on time-series forecasting, mostly with foundational models and edit-based methods. I am currently a Associate Research Fellow at Deakin University’s Applied AI Institute (A2I2).
Self-paced curriculum learning that combines cluster-based data reduction with a multi-armed bandit sampler, matching baselines while using up to 100× fewer training samples.
ACL 2026
An architecture-agnostic error-correction module that plugs into any forecaster without retraining; decomposes corrections into trend and seasonal components for robust gains across models and datasets.
ICML 2026
Guiding Reinforcement Fine-tuning with Intrinsic External Episodic Memory reward.
TMLR 2025
Non-parametric inference-time alignment with episodic memory; sample-efficient alignment under sparse feedback.
EMNLP 2025
Training-free token-level activation steering using episodic memory; adaptive alignment across safety & style.
ACL 2025
RL-based prompt example selection from episodic memory to boost generalization across NLP tasks.
ECAI 2024 (Oral)
Email: [email protected] · Phone: +61 412 242 886
Live view of where visitors are coming from.