Failure Detection for Surgical Robot Imitation Policies via Flow-Matching World Modeling
The paper introduces FoMo-FD, a failure-detection method for surgical robot imitation policies that trains an action-conditioned flow-matching world model on nominal short-horizon visual dynamics,...
1. Introduction: The Runtime Safeguard Challenge in Autonomous Surgery
Visuomotor imitation learning offers a promising route toward surgical robot autonomy by training policies on expert demonstrations gathered during teleoperated procedures. However, deploying these learned policies safely in real-world clinical settings remains a major challenge. Autonomous surgical manipulation demands exceptional precision within highly complex, dynamic, and variable environments. When an imitation policy encounters deployment distribution shifts—such as unmodeled anatomical variations, scene disturbances, or minor physical slippage—it can experience execution failures that threaten patient safety.
Developing reliable runtime failure monitors for autonomous surgery is complicated by three main obstacles:
- Data Scarcity: Negative execution examples and failure demonstrations are inherently rare, dangerous, and impractical to collect systematically in physical surgical environments.
- Environmental Variability: Surgical dynamics involve flexible, soft, and visually complex anatomical structures that lead to highly variable operational conditions.
- Workflow Sensitivity: Unreliable monitors that trigger frequent false alarms disrupt the surgeon’s operational flow and induce intervention fatigue.
To address these challenges, Zhefeng Huang et al. (2026) introduced FoMo-FD (Flow-Matching World Model for Failure Detection). FoMo-FD is an execution-level runtime monitor designed to detect fast, low-level visual-action inconsistencies during robot policy execution. Rather than requiring explicit failure dataset collection or assuming specific failure modes in advance, the system trains exclusively on successful execution data to learn nominal short-horizon visual-action dynamics.
2. Core Architecture: How FoMo-FD Monitors Visual-Action Consistency
FoMo-FD monitors robot execution by converting multi-view camera streams into compact latent representations, predicting nominal endpoint dynamics over short time windows, and measuring whether the realized execution deviates from expected behavior.
+-----------------------------------------------------------------------------------+
| FoMo-FD Pipeline |
+-----------------------------------------------------------------------------------+
| [ Multi-View Vision Input ] ---> ( Frozen DINOv2 Backbone ) |
| | |
| v |
| ( Feature-Space β-VAE ) |
| | |
| v |
| [ Intervening Action Chunk (A_{t-K}) ] ---> [ Visual History (Z_{t-K}) ] |
| | | |
| +----------------+---------------+ |
| | |
| v |
| ( Conditioning Context c_{t-K} ) |
| | |
| v |
| ( Flow-Matching World Model ) |
| | |
| v |
| ( Inverse-Transport Backward ODE Integration ) |
| | |
| v |
| [ View Nonconformity Score (r^v_t) vs Threshold (λ^v) ] |
+-----------------------------------------------------------------------------------+
The pipeline operates through four main structural stages:
-
Visual Latent Encoding:
- Multi-view visual observations (such as fixed workspace cameras, , and tool-mounted wrist cameras, ) are processed through a frozen DINOv2 vision foundation model backbone to extract high-level feature maps .
- To keep computational requirements manageable for real-time scoring, these features are compressed using an offline-trained feature-space -VAE ().
- The VAE incorporates a channel-utilization regularizer during training. This regularizer penalizes concentrated variance across channels, encouraging the feature representation to distribute information across the latent space rather than collapsing into a small subset of channels:
- After offline training, the VAE is frozen, and its deterministic posterior mode (where ) serves as the compact visual latent tensor.
-
Flow-Matching World Model Mechanism:
- Instead of recursively predicting frame-by-frame visual outputs—a process prone to error accumulation over extended rollout steps—the world model directly predicts the endpoint latent () of a short execution window.
- For a conditioning timestep, the model receives context , where represents a visual latent history window of length and represents the intervening commanded action chunk of length . The total monitored execution window spans timesteps.
- To fuse visual and action information, the architecture uses a dual action-conditioning mechanism: global action embeddings provide adaptive modulation across transformer layers, while action tokens feed directly into cross-attention mechanisms.
- During training, a vector field is learned to match target velocity along linear interpolation path , where and .
- Auxiliary Training Loss (): To regularize the visual-action representations, an auxiliary predictor head is added during training to predict the single-step next visual latent with loss . The combined training objective is (with and ). The auxiliary head is discarded after training and is not used for runtime inference.
-
Inverse-Transport Nonconformity Scoring:
- At test time, the system evaluates the visual-action consistency of an observation window ending at current timestep .
- The conditioning context is drawn from timesteps prior: , containing history and action chunk .
- The realized visual endpoint latent is integrated backward from to along the learned flow-matching ODE conditioned on :
- The view-specific nonconformity score is computed as the normalized squared distance of in the base Gaussian space:
- If physical interaction matches nominal dynamics, the inverse-transported latent lands near the base Gaussian origin, yielding a low score. If visual transitions violate nominal dynamics given the commanded actions, rises.
-
Conformal Threshold Calibration:
- To establish view-specific alarm thresholds () without failure data, the authors apply conformal prediction across successful calibration rollouts.
- For each calibration rollout and camera view , the peak nonconformity score is extracted: .
- Given a target miscoverage rate (e.g., ), scores are sorted (), and the task-specific threshold for view is calibrated as:
- At deployment, an alarm is triggered independently for camera view whenever .
3. Empirical Performance: Simulation and Real-World dVRK Experiments
Huang et al. evaluated FoMo-FD across four surgically relevant manipulation tasks:
- Needle Pickup (Simulation in NVIDIA Isaac Sim)
- Ring over Post (Simulation in NVIDIA Isaac Sim)
- Tissue Retraction (Physical da Vinci Research Kit / dVRK)
- Shunt Insertion (Physical da Vinci Research Kit / dVRK)
The evaluation covered 20 distinct failure modes across 320 failed policy rollouts and 80 successful policy rollouts. Failure types included grasp acquisition loss, spatial misalignments, lateral drift, blood occlusion, tissue disturbances, camera freezes, and action execution mismatches.
The table below summarizes quantitative performance comparing FoMo-FD (WM-NC) against visual feature prediction-error scoring (WM-PE) and observation-level visual anomaly baselines (logpZO and RND):
| Method / Input | Fixed-Camera FDR (%) | Fixed-Camera FAR (%) | Wrist-Camera FDR (%) | Wrist-Camera FAR (%) |
|---|---|---|---|---|
| WM-NC (FoMo-FD) | 45.9% | 10.0% | 96.6% | 1.3% |
| WM-PE (Prediction Error) | 16.9% | 3.8% | 52.2% | 7.5% |
| logpZO (FAILDetect variant) | 34.4% | 0.0% | 45.3% | 3.8% |
| RND (FIPER variant) | 37.8% | 0.0% | 42.8% | 1.3% |
FDR = Failure Detection Rate (higher is better); FAR = False Alarm Rate (lower is better).
- Dominance of Wrist-View Monitoring: Wrist-camera monitoring using inverse-transport scoring (WM-NC) achieved the highest overall detection rate (96.6% FDR at a 1.3% false alarm rate).
- Window Dynamics vs. Single-Frame Anomalies: Observation-level anomaly detectors (logpZO and RND) struggled on complex manipulation tasks like Shunt Insertion, achieving lower overall wrist FDRs (45.3% and 42.8%, respectively). Low-level execution failures (e.g., subtle instrument slippage or unexpected tissue resistance) rarely produce out-of-distribution individual frames. Instead, they manifest as temporal mismatches between commanded actions and visual responses over multi-step windows.
- Inverse Transport vs. Prediction Error: Scoring inverse-transport nonconformity in base Gaussian space (WM-NC) significantly outperformed visual latent prediction error (WM-PE, 52.2% wrist FDR). Integrating realized endpoints backward through the flow ODE captures dynamic trajectory nonconformity far more reliably than feature-space Euclidean distance metrics.
- Real-Time Computational Throughput: Benchmark testing on a desktop workstation (Intel Core Ultra 9 285K CPU, NVIDIA RTX 5090 GPU) confirmed that the full scoring pipeline—including DINOv2 visual encoding, VAE compression, and backward ODE integration—processes scoring windows at 13.98 Hz. Because the physical dVRK control loop operates at 10 Hz, this throughput validates that FoMo-FD is computationally suitable for real-time online deployment.
4. Deep-Dive: Key Architectural Ablations
To isolate the structural factors driving FoMo-FD’s monitoring performance, the authors conducted systematically controlled ablation experiments across camera configurations, conditioning pathways, and prediction horizons.
Action Conditioning Pathways
The authors evaluated four distinct action-conditioning modes within the flow-matching architecture:
- No Action Conditioning (NoAct)
- Action-Token Cross-Attention Only (Tok)
- Global Action Modulation Only (Mod)
- Combined Modulation and Cross-Attention (Mod+Tok)
Wrist-Camera FDR by Action Conditioning Mode:
-------------------------------------------------------
No Action Conditioning (NoAct) : [======== ] 87.2%
Cross-Attention Tokens (Tok) : [========= ] 88.8%
Global Modulation Only (Mod) : [========== ] 90.6%
Combined (Mod + Tok) : [=========== ] 96.6%
-------------------------------------------------------
Combining global action modulation with local token cross-attention enables the flow-matching transformer to track physical control commands precisely, producing a 9.4% absolute gain in wrist-view failure detection over unconditioned visual dynamics models.
Endpoint Horizon Length ()
The length of the endpoint prediction horizon () proved critical to detection sensitivity across total window horizons (with ):
Wrist-Camera FDR across Endpoint Horizons (K):
-------------------------------------------------------
K = 1 step (H = 5) : [==== ] 42.2%
K = 2 steps (H = 6) : [====== ] 59.8%
K = 3 steps (H = 7) : [========= ] 81.3%
K = 4 steps (H = 8) : [=========== ] 96.6%
-------------------------------------------------------
Single-step predictions () frequently fail to flag execution anomalies because single-frame transitions diverge minimally from nominal dynamics. Extending the prediction horizon to steps () allows subtle visual-action mismatches to accumulate over time, yielding a strong nonconformity signal without elevating false alarm rates.
Camera Placement Geometry
Across all evaluated tasks, wrist-mounted cameras consistently outperformed fixed workspace cameras. In surgical manipulation, key interactions—such as needle grasping, tissue tensioning, and vessel shunting—take place in small regions near the tool tip. Local wrist-mounted cameras capture these fine-grained instrument-tissue dynamics directly, whereas fixed global cameras are vulnerable to workspace clutter, changing background lighting, and distance attenuation.
5. Stated Limitations and Structural Scope
While FoMo-FD provides an effective execution-level safeguard, the authors highlighted two main operational limitations:
- Calibration Rollout Dependency: Although FoMo-FD eliminates the need for failure dataset collection, threshold setting via conformal prediction still requires a small set of successful policy rollouts ( in these experiments) under the target deployment policy distribution to maintain exchangeability.
- Thresholding Mechanics: The current implementation uses an episode-level peak score threshold () rather than a progress-conditioned or time-varying threshold band. Because surgical tasks vary in execution tempo across non-aligned rollouts, applying fixed global thresholds can introduce sensitivity to temporal phase variations.
System Role in Safety Architectures
In broader surgical AI safety architectures, FoMo-FD functions specifically as an execution-level monitor. It verifies whether local physical dynamics match expected visual-action trajectories.
+-------------------------------------------------------------------------------+
| Layered Surgical AI Safety Architecture |
+-------------------------------------------------------------------------------+
| High-Level Semantic Monitors (VLMs / Task Progress Monitors) |
| - Tracks task progression, semantic goals, procedural logic. |
+-------------------------------------------------------------------------------+
| Execution-Level Runtime Monitor (FoMo-FD) |
| - Tracks short-horizon visual-action consistency and physical interaction. |
+-------------------------------------------------------------------------------+
High-level monitors (e.g., Vision-Language Models or state-progress graphs) evaluate long-horizon semantic goals (whether the policy is performing the correct procedural step), whereas FoMo-FD provides continuous verification of physical interaction dynamics (whether active tool manipulation is unfolding safely).
6. Key Takeaways for AI Safety & Robotics Practitioners
- Nominal-Only Training Works: Runtime monitors do not require failure demonstrations to detect execution anomalies if they model action-conditioned visual dynamics effectively.
- Multi-Step Windows Outperform Single Frames: Scoring multi-step window transitions () via flow inverse-transport provides a substantially stronger failure signal than single-frame observation anomaly detection or raw prediction error.
- Tool-Centric Sensing is Essential: Tool-mounted local views (like wrist cameras) capture crucial fine-grained instrument-tissue interactions that are frequently obscured or diluted in fixed workspace camera views.
Read the full paper on arXiv · PDF