Last modified: 2026-06-09
Abstract
As virtual workspace operations grow increasingly complex, balancing automated efficiency with human expert decision-making has become a critical challenge. Traditional automation lacks the flexibility to handle dynamic UI changes, while continuous manual operation leads to significant cognitive fatigue.
This research proposes a Shared Autonomy framework centered on the Action Chunking Transformer (ACT) to facilitate seamless human-AI collaboration. The system is designed to integrate DINOv3[1] as a frozen vision backbone to extract high-fidelity semantic and geometric features, providing the model with a robust understanding of complex virtual layouts. By leveraging the ACT architecture, the framework aims to generate smooth, multi-step action sequences derived from expert demonstrations.
A key innovation of this work is the planned introduction of an uncertainty-aware gating mechanism. By monitoring the predictive variance of the action head, the system is designed to autonomously adjust its control authority. When encountering ambiguous scenarios, the framework will proactively trigger a control handover to the human user, ensuring both operational safety and precision.
The anticipated experimental evaluations in virtual archiving tasks are expected to demonstrate that the proposed method can significantly reduce the intervention frequency and human workload. This study seeks to provide a robust and scalable solution for future shared-control systems by combining imitation learning with realtime uncertainty estimation.