Semi-autonomous robotic disassembly is not only a question of making an arm move correctly: the operator must also understand what the robot sees, what it intends to do and when intervention is needed. SARDiM brings perception, assisted manipulation and mixed-reality visualisation into the same battery-disassembly framework, while evaluating control assistance separately from the choice of display.
The Problem: Controlling a Robot Through an Incomplete View
A direct view of a workcell provides natural depth cues, but remote work often replaces that view with a camera feed. A flat image can make the relationship between a tool and a component difficult to judge, especially when the robot obscures the target. Mixed reality can restore some spatial context by presenting a reconstructed scene, but a richer display is not automatically an easier or safer interface.
The paper therefore studies two related questions: whether semi-autonomous assistance changes manipulation performance, and how different views affect the operator’s interaction with that assistance. Keeping these questions distinct is important because the display communicates information, whereas the control mode determines how human commands become robot movements.
From Segmented Components to Assisted Motion
FastSAM supplies visual segments from which the system estimates component geometry. Cluster analysis contributes component centroids and orientations; size and disassembly priority help organise the task representation. A centroid is a useful reference for selecting a component, but is not by itself a collision-free trajectory or a validated tool-contact pose.
MoveIt provides the motion-planning component for the Franka arm. Joint impedance control and force feedback support interaction beyond simply commanding a sequence of positions. Impedance describes how the arm responds to displacement and contact: the robot’s mechanical response matters when the target position and the physical component do not align perfectly.
ROS, Unity and MATLAB connect robot communication, the mixed-reality interface and control-related processing. Their roles are complementary. A scene shown in Unity must agree with the robot’s coordinate frames, and a geometrically plausible movement must still respect the arm’s limits and surrounding obstacles.
graph TD
A[Camera observations] --> B[FastSAM segments and component geometry]
B --> C[Task selection and motion planning]
D[Operator commands] --> C
C --> E[Franka arm and impedance control]
B --> F[Direct, monitor or mixed-reality view]
E --> F
F --> D
Explanatory diagram created for this article, not a reproduced paper figure: perception, control and visualisation form a feedback loop rather than a one-way instruction display.
What the Comparison Actually Measures
The study combines two control modes with four interface methods. The modes are manual control and semi-autonomous variable autonomy. The views are direct observation, a monitor feed, mixed reality with a monitor feed, and point-cloud mixed reality. Together they produce eight conditions.
This structure makes within-interface comparisons especially useful: manual and assisted operation through the same view differ primarily in the control condition. Conversely, comparing displays within one control mode helps examine the interface. A comparison that changes both at once measures the combined system, not an isolated effect of mixed reality.
The semi-autonomous mode reported 40.61% fewer joint-limit violations. The semi-autonomous point-cloud MR condition completed tasks 2.33% faster than manual direct-view control. The larger change concerns limit violations, not speed; presenting the study as a dramatic cycle-time improvement would miss its more relevant finding.
Interpretation and Practical Boundaries
Fewer joint-limit violations indicate that assistance can help operators avoid problematic arm configurations in the tested tasks. They do not establish that the entire system is safe: collision risks, registration errors, contact forces and electrical hazards are different properties that need their own evidence.
Likewise, the reported time comparison does not establish reduced technician training time, improved anomaly recovery or performance across every battery design. The contribution is a modular framework and a structured comparison of assistance and visual feedback. Repeating the original experiment would require matched task instructions, scene and robot geometry, calibration, perception settings, controller parameters and scoring rules; the publication provides the reference for those choices, rather than this overview implying an exact turnkey reproduction.
Testing and Real-World Use
This could be tested in a simulated or suitably supervised workcell by comparing manual and assisted control through the same display, recording task time, joint-limit events and intervention needs. In real maintenance and disassembly, the approach could help a remote operator understand component geometry and supervise repetitive motion while retaining control over uncertain steps.
Paper and Figures
The conference paper was published on 21 August 2024, after the June conference. The open-access author-institution copy contains the original system and experiment figures; the university record identifies its publication details.