Publications

Papers, collaborations, and evolving research directions.

This page keeps public-facing publication information concise. Works that are still under review use venue-neutral status labels until their outcomes can be shared.

Findings of EMNLP 2026 First Author 2026

ASIL: Replacing Screenshot-and-Click with Structured State and Semantic Actions

Rui Xie, Lu Chen

An agent-native interface that replaces screenshot-and-click control with structured software state and code-executable semantic actions, reaching 81.6 on GPT-5.4 under a 15-action budget versus 6.6 for screenshot GUI control under 50 steps.

arXiv Technical Report, 2026 Core Contributor 2026

Qwen-CUA: Native Computer Use for (almost) Everything

Qwen Team & XLang Lab

A technical report presenting a native computer-use agent that operates software from screenshots through keyboard and mouse actions, supported by large-scale verifiable tasks and interactive rollouts.

ECCV 2026 First Author 2026

GUIDE: Resolving Domain Bias in GUI Agents through Real-Time Web Video Retrieval and Plug-and-Play Annotation

Rui Xie, Zhi Gao, Chenrui Shi, Zirui Shang, Lu Chen, Qing Li

A training-free, plug-and-play framework that retrieves task-relevant tutorial videos and distills transferable planning and grounding knowledge for domain-specific GUI agents, yielding +4.47 to +7.48 point gains on OSWorld without modifying agent parameters.

Under Review Co-first Author 2026

MatToolBench: A Real-Environment Benchmark for Evaluating Multimodal Agents on Professional Materials Science Software

Mei Wu, Rui Xie, Runyu Zhang, Lu Chen, Bo Chen, Kai Yu, Xin Chen

A real-environment benchmark that evaluates multimodal agents on professional materials science workflows spanning GUI tools, code execution, and cross-tool coordination.

Under Review Third Author 2025

Mobile-Env: Building Qualified Evaluation Benchmarks for LLM-GUI Interaction

Danyang Zhang, Zhennan Shen, Rui Xie, Situo Zhang, Tianbao Xie, Zihan Zhao, Siyuan Chen, Lu Chen, Hongshen Xu, Ruisheng Cao, Kai Yu

A benchmark effort for evaluating LLM-based GUI interaction in mobile environments with isolated tasks, simulator support, and behavior analysis.

AAAI 2026 Contributing Author 2025

TongUI: Internet-Scale Trajectories from Multimodal Web Tutorials for Generalized GUI Agents

Bofei Zhang, Zirui Shang, Zhi Gao, Wang Zhang, Rui Xie, Xiaojian Ma, Tao Yuan, Xinxiao Wu, Song-Chun Zhu, Qing Li

A large-scale data construction effort that turns multimodal web tutorials into GUI trajectories for generalized agent training and evaluation.