My work centers on spatio-temporal intelligence, time-series forecasting, and multimodal learning, with a growing focus on embodied foundation models and AI infrastructure.
15
Publication records
9
Service venues
5
Research tracks
Selected research output
Publications
Peer-reviewed papers and selected preprints, with first-author and featured work prioritized.
TS-Memory: Plug-and-Play Memory for Time Series Foundation Models
PinnedFeatured
We propose TS-Memory, a lightweight plug-and-play memory adapter that augments frozen Time Series Foundation Models. TS-Memory distills offline retrieval's distributional knowledge into a compact neural module via Parametric Memory Distillation, enabling retrieval-free deployment with constant-time inference while achieving state-of-the-art performance across diverse benchmarks.
Visual Reasoning over Time Series via Multi-Agent System
PinnedUnder ReviewFeatured
Weilin Ruan,
Yuxuan Liang
Under Review · 2026
We propose MAS4TS, a tool-driven multi-agent framework for general time-series tasks under an Analyzer-Reasoner-Executor paradigm. It integrates visual reasoning over time-series plots, latent trajectory reconstruction, and gated inter-agent communication to improve cross-task generalization and inference efficiency.
We propose OccamVTS, a novel framework that distills large vision models to only 1% of their original parameters for efficient time series forecasting, demonstrating that extreme parameter reduction can be achieved while maintaining strong predictive performance through innovative distillation techniques.
We propose RAST, a universal framework that integrates retrieval-augmented mechanisms with spatio-temporal modeling to address limited contextual capacity and low predictability in traffic prediction. Our framework consists of three key designs: Decoupled Encoder and Query Generator, Spatio-temporal Retrieval Store and Retrievers, and Universal Backbone Predictor that flexibly accommodates pre-trained STGNNs or simple MLP predictors.
We propose ViST, a novel vision-driven framework that transforms raw spatio-temporal data directly into a low-dimensional visual space for more effective and scalable forecasting. Our method features Multi-view Vision Transformation, Multi-modal Conditional Reconstruction, and Efficient Cross-modal Fusion Mechanism, achieving state-of-the-art accuracy with significantly reduced computational cost.
This paper proposes a novel spatio-temporal unitized model for traffic flow forecasting that effectively captures complex dependencies across both space and time dimensions, achieving state-of-the-art performance on multiple benchmark datasets.
We propose MMLoad, a novel diffusion-based multimodal framework for multi-scenario building load forecasting with three innovations: Multimodal Data Enhancement Pipeline, Cross-modal Relation Encoder, and Scenario-Conditioned Diffusion Generator with uncertainty quantification, establishing a new paradigm for multimodal learning in smart energy systems.
MultimodalEnergy ForecastingSmart Buildings
Low-rank Adaptation for Spatio-Temporal Forecasting
PinnedFeaturedOral
This paper presents ST-LoRA, a novel low-rank adaptation framework as an off-the-shelf plugin for existing spatial-temporal prediction models, which alleviates node heterogeneity problems through node-level adjustments while minimally increasing parameters and training time.
We present DeepUHI, a heat equation-based framework that models urban heat island effects through thermodynamic cycles and thermal flows, integrating multimodal environmental data to achieve precise street-level temperature forecasting, now deployed as a real-time warning system in Seoul.
This paper proposes Time-VLM, a novel multimodal framework that leverages pre-trained Vision-Language Models (VLMs) to bridge temporal, visual, and textual modalities for enhanced time series forecasting.
We investigate the anchoring effect in Large Language Models, exploring whether LLMs are affected by anchoring bias, the underlying mechanisms, and potential mitigation strategies. We introduce SynAnchors dataset and show that LLMs' anchoring bias exists commonly with shallow-layer acting and is not eliminated by conventional strategies, while reasoning can offer some mitigation.
PaperNatural Language ProcessingCognitive BiasLarge Language Models
VQualA 2025 Challenge on Face Image Quality Assessment: Methods and Results
Featured
We introduce the VQualA 2025 Challenge on Face Image Quality Assessment, part of ICCV 2025 Workshops. Participants developed efficient models (<0.5 GFLOPs, <5M parameters) predicting Mean Opinion Scores (MOS) under realistic degradations. The challenge attracted 127 participants, resulting in 1519 valid final submissions, contributing to practical FIQA solutions.
LDM4TS: Latent Diffusion Model for Time Series Forecasting
Featured
Weilin Ruan,
Siru Zhong,
Haomin Wen,
Yuxuan Liang
Under review · 2025
This paper introduces LDM4TS, a novel latent diffusion model for time series forecasting that transforms time series into multiple image representations and leverages diffusion models to enhance forecasting capabilities.
We introduce GSTRL, a novel game-theoretic reinforcement learning framework that addresses collaborative public resource allocation by modeling it as a cooperative potential game and incorporating spatio-temporal learning to capture crowd dynamics, outperforming existing methods on real-world datasets.
We present a novel approach that combines Gaussian Splatting with kinematic knowledge to reconstruct and simulate articulated objects, enhancing both geometric detail and articulation dynamics while overcoming the limitations of previous implicit-based models.