Hello, I'm Jiashu Yang (杨佳澍). My research interests focus on open-world visual understanding and large-scale foundation models (language, vision, and action), with a commitment to developing practical and deployable intelligent algorithms. In the realm of LLMs, I lead the open-source community Wenyuan Pavilion, dedicated to Chinese-culture-centered language models. Since 2023, I have been pursuing Computer Vision research supervised by Yian Zhao and Chaoran Feng. Additionally, I have been exploring LLMs mentored by Xuxin Cheng since 2024.
If you are interested in discussing or collaborating with me, please feel free to contact me via email.
📝 Publications
- EvaGaussians: Event Stream Assisted Gaussian Splatting from Blurry Images
- Kongzi: A Historical Large Language Model with Fact Enhancement
- Tune-Your-Style: Intensity-tunable 3D Style Transfer with Gaussian Splatting
- Breaking the Vicious Cycle: Coherent 3D-GS from Sparse & Blurred Views
- Look, Zoom, Understand: The Robotic Eyeball for Embodied Perception
- AutothinkRAG: Complexity-Aware Control of Retrieval-Augmented Reasoning for Image-Text Interaction
📖 Education
-
- Sincerely looking for PhD positions for fall admission!
💻 Projects
-
LongHorizonOS: The OS for Long-Running Agents
An OS-like exoskeleton that wraps existing agent harnesses with time-sliced execution, investment-gated restarts, watchdog stop-loss, and versioned progress graphs. Evaluated on the 46-task LongHorizonBenchmark (LHTB), achieving ~2.4× fewer input tokens with 37/42 reward parity or better.
-
Document Intelligence Survey
Interactive web demo for the survey paper "Beyond OCR: A Survey of Document Intelligence from Structured Perception to Multimodal Question Answering". Provides a structured overview of document understanding techniques spanning OCR, layout analysis, information extraction, and multimodal QA.
-
DocThinker: All-in-One RAG SystemDocThinker is a next-generation intelligent agent system designed to break the limitations of traditional RAG. It builds a structured, brain-like memory system using Knowledge Graphs.
-
Wenyuan Pavilion: Ancient Chinese Language CommunityI lead a research community focused on domain-specific models for Ancient Chinese and Literature. We have open-sourced various models and datasets.
💼 Experience
-
2026.06 - Present Qianlima Talent Program, Qianli Technology (Afari)
World models and VLA, including representation learning, data construction, model post-training, and RL frameworks.I lead a team of PhD research interns developing next-generation vision-language-action (VLA) architectures: Hefei Huang, Jianyu Zhang, and Zhiyuan Liang.
I welcome talented and motivated researchers to join our team and help shape the next generation of VLA models. Feel free to reach out!
-
2025.12 - 2026.05 Research Intern, Meituan, Beijing
Longcat Interaction Team -
2025.07 - 2025.09 Research Intern, Shanghai Jiao Tong University
School of Artificial Intelligence -
2025.04 - 2025.07 Research Intern, ByteDance, Beijing
Applications of Large Language Models -
2023.11 - 2024.08 Research Intern, CASIA
Institute of Automation, Chinese Academy of Sciences
🎖 Honors
- 2023: Robocup, Advanced Vision Track - National First Prize
- 2024: China University Computer Competition - National Third Prize
- 2024: Robocup, Advanced Vision Track - National Third Prize
About Life
- Table Tennis — Regular player, enjoy the fast pace and tactical depth of the sport.
- Martial Arts — Practitioner, drawn to the discipline, tradition, and physical & mental cultivation.