Curriculum Vitae

I am an undergraduate at Tsinghua University pursuing dual degrees in Computer Science and Technology and Economics and Finance, with graduation expected in July 2027. My experience spans data synthesis, experimental design, evaluation, and paper writing, alongside research internships at ZTE and Meituan.

My research interests are reinforcement learning for models, inference optimization, and trajectory analysis. I work on visual perception, probabilistic memory for language models, and demonstration-guided GUI agents.

01 Research Interests

Reinforcement Learning

Model post-training · Visual curricula

Inference Optimization

Efficient memory · Device-cloud inference

Trajectory Analysis

GUI demonstrations · World models

02 Education

Sep. 2023Jul. 2027Expected

Tsinghua University Undergraduate

B.Eng. in Computer Science and Technology · B.Econ. in Economics and Finance (dual degree)

GPA 3.3 / 4.0

Software Engineering Computer Graphics Game Theory Student Research Training Machine Learning
Sep. 2020Jun. 2023

The High School Affiliated to Northwestern Polytechnical University

High school

03 Publications & Patents

NAACL 2027 Under review Feb. 2026 – Present

Boosting Fine-Grained Visual Perception via Multi Granularity Data Synthesis and Local-to-Global Curriculum Learning

W. Yan#, Y. Liu#, D. Zou et al.

Vision-language models can describe a scene while still missing object attributes or spatial relationships. This work addresses that gap with a multi-granularity data pipeline and Local-to-Global Curriculum Reinforcement Learning (LoGo-CRL).

The pipeline combines object tags and grounded regions with teacher-generated questions and answers, then filters examples through agreement between two teacher models and the student model's current capabilities. Training progresses from object grounding and attribute recognition to relationships between objects, then to counting and detailed scene understanding. Each stage uses task-specific rewards so that local perceptual skills support later global reasoning.

NeurIPS 2026 Under review Mar. 2026 – Aug. 2026

From Cache to Belief: Online Probabilistic Memory for Long-Context Language Models

J. Wu, Y. Liu & Y. Yue* et al.

BeliefMem studies how a language model can retain useful evidence beyond its local attention window while representing uncertainty in compressed history. It adds a lightweight probabilistic memory adapter to a frozen language model.

A fixed set of Bayesian memory slots is updated online. The model reads a query-dependent belief from those slots and combines it with local evidence through a product of Gaussian experts. Shared responsibilities govern both reading and writing, keeping retrieval aligned with memory updates. Experiments on six LongBench tasks evaluate long-context reasoning, while ablations and slot interventions examine whether gains come from distributed content updates rather than a larger local window.

NAACL 2027 Under review Sep. 2025 – Mar. 2026

You Only Show Once: Robust GUI Automation via One-Shot Human Demonstration

Y. Huang#, Y. Liu#, W. Yan et al.

YOSO-Agent uses a single human demonstration to guide GUI task execution. It preserves action-aligned screenshots and visual references, giving the agent concrete evidence about what to do and where to act when the interface differs from the demonstration.

A Task Manager checks progress against the reference trajectory; an Action Agent grounds the next interaction in the current screen; and a Reflection Agent diagnoses deviations and supports recovery. The grounding model is trained with varied visual appearances. Evaluation covers both target localization under interface perturbations and end-to-end task completion. The framework follows a supplied demonstration for the task; selecting the appropriate demonstration is outside its scope.

PPSN 2026 Accepted Aug. 2025 – Dec. 2025

A Scalable Benchmark Test Suite for Dynamic Multi-Objective Optimization with a Changing Number of Objectives

K. Shang, Z. Xiao, Y. Liu et al.

Dynamic optimization benchmarks often change the objective functions themselves when the number of objectives changes. This makes it difficult to isolate an algorithm's ability to adapt to adding or removing evaluation criteria.

The benchmark defines a problem with a fixed maximum set of objectives, then activates different subsets over time. The underlying functions remain unchanged. Minus-DTLZ and Minus-WFG formulations avoid degenerate Pareto fronts that can arise when selecting subsets of conventional benchmark objectives. The resulting suite supports controlled comparisons of algorithms as objectives are added or removed.

Read the paper on arXiv →

Chinese invention patent Under substantive examination App. No. 2026105681825

A Device-Cloud Collaborative Multimodal Inference Method and System

Y. Liu

This method dynamically allocates perception and reasoning between an on-device model and cloud computation. The system targets the trade-off between local inference cost and the need to protect sensitive information during multimodal processing.

04 Research Projects

Jun. 2024Jun. 2025

Graph-Generating LLMs for Drug Discovery

Student Research Training (SRT), Department of Computer Science and Technology

I built training-data and prompt synthesis pipelines, applied reinforcement learning to post-train DeepSeek-R1 and other language models, and compared their graph-generation capabilities. I also proposed a benchmark for evaluating whether generated graphs satisfy explicit rules and constraints, motivated by graph-based applications in drug discovery.

05 Internships

Jun. 2026Jul. 2026

Meituan

World Model Intern · Autonomous Vehicles

I processed multimodal data and trained world models for autonomous driving, then evaluated and optimized perception modules for robustness in complex environments.

Aug. 2025Mar. 2026

ZTE Corporation

AI Research Intern · Multimodal LLM Inference

I researched inference acceleration and visual perception for multimodal language models. I independently built data synthesis and evaluation pipelines to support fine-tuning and repeated ablation studies.

06 Service & Leadership

Jun. 2025Present

Student Association for Science and Technology, Department of CST

Vice President · AI Agent Division

I led development of the official game player and visualizer for the 29th AI Agent Contest, used by participants across Tsinghua University. I also coordinated university-wide contests, including problem design, scheduling, and technical support.

I participated in research fieldwork in Japan and volunteer teaching in Qinghai, with leadership responsibilities in communications and teaching activities.

07 Honors & Awards

  • Outstanding Social Work Scholarship, Tsinghua University2025
  • Third Prize, Youth Digital Innovation Competition (Beijing Xicheng)2025
  • Winning Prize, 7th “Bambu Lab Cup” Software Design Competition, Department of EE2024

08 Skills

Languages
PythonC / C++C#Rust
Deep Learning
PyTorchTransformersLLM trainingFine-tuningEvaluation
Engineering
LinuxMulti-GPU trainingExperiment managementGitCode review
Full Stack
Wordle web game (Yew / WebAssembly) and the yaoJ online judge framework (actix-web). View on GitHub
Communication
Mandarin Chinese (native); English (academic reading, writing, and paper preparation).