arXiv · 2026
TeleBoost explains why high-quality video alignment depends on staged supervised shaping, reward-driven reinforcement learning, preference refinement, diagnostics, and systems engineering working together.
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2026) · 2026 · pp. 27260-27269
A practical explanation of why sequence-level rewards are too coarse for visual generation, and how Visual Preference Policy Optimization turns perceptual structure into localized learning signals.
European Conference on Computer Vision (ECCV 2026, accepted) · 2026 · forthcoming
TaRoS examines how fixed reward components can become optimization targets rather than useful feedback, then reshapes video GRPO signals by component, group sparsity, and saturation.
ACM Multimedia 2026 (accepted) · 2026 · forthcoming
OTCA reframes diffusion-model alignment as process credit assignment, decomposing an outcome reward over a denoising trajectory and allocating multiple objectives where they are most useful.
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2026) · 2026 · pp. 34408-34417
An intuitive and technical guide to Bayesian Prior-Guided Optimization, which treats reward confidence as part of visual policy optimization instead of trusting every group comparison equally.
SIGGRAPH Asia 2025 Conference Papers · 2025 · pp. 1-11
Uni-Inter maps three previously separate 3D interaction settings into a shared volumetric language, then predicts probabilistic human motion while preserving each task’s entities and constraints.
IEEE Transactions on Neural Networks and Learning Systems · 2024 · vol. 35(5) · pp. 7048-7059
Moderate Hard Example Modulation emphasizes difficult fine-grained samples while bounding their relative influence so that noise and memorized exceptions do not dominate training.
arXiv · 2024
AntEval evaluates LLM-driven agents at the process level: whether information actually moves between participants and whether intent is expressed in a human-like way.
IEEE Transactions on Multimedia · 2024 · vol. 26 · pp. 4389-4400
IcoCap combines image semantics with video features and selects captions that match the resulting visual content instead of preserving noisy original labels.
IEEE/CVF International Conference on Computer Vision (ICCV 2023) · 2023 · pp. 217-227
MAAL studies low-sample affordance learning by combining 3D object and robot modalities, retaining successful actions, and scoring candidates through reconstruction rather than appearance alone.
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2022) · 2022 · pp. 10463-10472
SEEG separates rhythmic and semantic cues, constrains generated motion with semantic information, and shows why gesture evaluation needs content-aware data and metrics.
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2022) · 2022 · pp. 9549-9559
Episodic Linear Probe repeatedly resets a simple classifier during training to distinguish what the classifier remembers from what the current representation makes easy to learn.
IEEE/CVF International Conference on Computer Vision (ICCV 2019) · 2019 · pp. 10402-10411
VrR-VG uses a counterfactual dataset test: remove image pixels, measure which relationships remain predictable from object labels and geometry, and refocus learning on visual evidence.