Resilience Evaluationfor Embodied Agents System

Measure how agents recover, remain stable, and sustain service as execution stress increases.We reveal the embodied tasks execution resilience in EAS process.

Habitat-Sim Explore stage showing the robot and human surveying the bedroom scene.
REPRESENTATIVE HABITAT-SIM EXECUTION

METRICS · 01

Three views of EAS resilience

Rebound measures recovery burden. Stability compares clean and stressed trajectories. GE finds sustained capacity across a complete stress grid.

Conceptual illustration of the full embodied-agent execution process, with the Rebound zone emphasized: stress interrupts object handling and a diagnosis, perception, and replanning loop restores execution.
STATISTICAL UNITRollout × closed window \([t_d,t_r]\)↓ \(C_{\mathrm{rec}}\)
METRIC LOGICClosed-window recovery burden
Interactive construction of Rebound recovery cost Recovery intensity is accumulated between disturbance time t d and recovery time t r, then normalized by remaining clean-reference work. t_d G_rec,j = weighted excess work ÷ clean remaining-work reference W*rem Σ CLOSED WINDOWS C_rec ↓ t_r

We aggregate normalized recovery burden within the closed window.

\(C_{\mathrm{rec}}\)lower is better

Recovery cost

Implemented estimator\[ \begin{aligned} I_{\mathrm{rec}}(t)&=\sqrt{\frac{g_{\mathrm{cog}}^2(t)+g_{\mathrm{phy}}^2(t)+g_{\mathrm{debt}}^2(t)}{3}},\\ G_{\mathrm{rec},j}&=\sum_{t\in\mathcal W_j} \frac{\max(\Delta\tau_{\mathrm{sim},t},1)}{\max(\tau_t^*,\epsilon)}I_{\mathrm{rec}}(t),\\ C_{\mathrm{rec},j}&=\frac{G_{\mathrm{rec},j}}{W_{\mathrm{rem}}^*(e,s_{d,j})+\epsilon},\\ C_{\mathrm{rec}}&=\sum_{j\in\mathcal J_{\mathrm{closed}}}C_{\mathrm{rec},j}. \end{aligned} \]
Definition
Normalized excess work required to restore acceptable execution.
Statistical unit
One rollout, within each closed window \([t_d,t_r]\).
Open Rebound

BENCHMARK · 02

Performance beyond task success

Ten EAS methods across 400 household tasks reveal distinct process-level strengths; no method dominates every resilience aspect.

LOWEST \(C_{\mathrm{rec}}\)6.5

ReAct · SR 38.5%

LOWEST \(\beta\)0.149

InnerMono · completion 67.5%

HIGHEST \(\lambda^*\)0.85

CoPAL · completion 72.6%

GLOBAL RESULTNo dominance in 10 EAS

Each method exhibits a distinct resilience signature.

TABLE 2 · INTERACTIVE

Compare all methods

Table 2 Exact benchmark values
Table 2. Mean ± reported variation across 400 household tasks. Arrows indicate preferred direction.
Method\(C_{\mathrm{rec}}\) ↓\(\beta\) ↓Stress capacity \(\lambda^*\) ↑SRCompletionSim. steps
ReAct6.5 ± 0.460.239 ± 0.0110.50 ± 0.04238.5% ± 3.1%57.8% ± 4.6%3,374 ± 285
Reflexion57.3 ± 5.070.299 ± 0.0210.66 ± 0.03556.3% ± 4.2%75.0% ± 3.8%4,390 ± 312
CycleVLA35.5 ± 2.910.400 ± 0.0260.72 ± 0.04722.2% ± 2.8%47.5% ± 4.1%4,688 ± 420
CLARE21.6 ± 1.680.344 ± 0.0160.72 ± 0.05228.4% ± 3.5%46.0% ± 4.5%4,249 ± 360
SayCan38.5 ± 2.490.283 ± 0.0130.78 ± 0.04657.0% ± 4.1%72.6% ± 3.7%4,682 ± 380
AgentEvolver19.2 ± 1.240.264 ± 0.0120.21 ± 0.01548.8% ± 3.9%66.8% ± 4.2%3,084 ± 250
InnerMono43.5 ± 2.690.149 ± 0.0070.53 ± 0.02649.5% ± 4.0%67.5% ± 4.3%3,670 ± 290
CoPAL11.9 ± 1.020.302 ± 0.0170.85 ± 0.07855.8% ± 4.5%72.6% ± 3.9%4,846 ± 410
SMART-LLM28.2 ± 2.200.339 ± 0.0180.40 ± 0.03655.7% ± 3.8%74.3% ± 4.1%5,085 ± 395
Baseline25.0 ± 2.030.200 ± 0.0090.79 ± 0.04253.0% ± 3.6%69.7% ± 3.5%3,671 ± 275

OPTIMIZATION · 03

Targeted resilience improve

Each intervention addresses one diagnosed execution failure while reporting both the intended gain and its cross-aspect trade-offs.

01 · REBOUNDFeedback-guided recovery
Crec59.11 → 33.73−42.94%
Recovery windows
3.13 → 2.38 −23.96%
Trade-off
GE stress-test task completion 0.887 → 0.819

We use World-Graph evidence and execution feedback to produce bounded recovery guidance.

02 · STABILITYConsistency state record
βstep0.3185 → 0.2552−19.87%
Paired β
0.6788 → 0.6772 −0.24%
Trade-off
Recovery windows 2.63 → 3.00

We use a read-only phase record to constrain replanning drift while preserving adaptation.

03 · GEBoundary long-tail control
Task completion · GE stress test0.833 → 0.917+10.08%
Formal GE
Incomplete → Complete
Trade-off
Crec 36.70 → 99.59

We use a bounded cognitive reset to convert persistent waiting into measurable boundary execution.

Table 3 Exact metric-guided optimization results
Targeted improvements under matched evaluation conditions. Percentage changes are relative to each target-specific baseline.
TargetVariantRebound CrecRecovery windows ↓βstepβ ↓GE completion ↑Formal GE
ReboundBaseline59.113.130.28810.68340.887
Optimized33.73 ↓42.94%2.38 ↓23.96%0.2079 ↓27.84%0.6870 ↑0.53%0.819 ↓7.67%
StabilityBaseline51.012.630.31850.67880.903
Optimized43.01 ↓15.68%3.00 ↑14.07%0.2552 ↓19.87%0.6772 ↓0.24%0.750 ↓16.94%
GEBaseline36.701.331.45940.69170.833Incomplete
Optimized99.59 ↑171.36%2.25 ↑69.17%0.7215 ↓50.56%0.6371 ↓7.89%0.917 ↑10.08%Complete
† Formal GE is reported only for GE stress tests because it requires a stress-response curve over the λ grid.

SIMULATION EXECUTION

Habitat-3.0 Simulation

Episode 75 connects instructions, plans, simulator feedback, metrics, and verification across six narrated acts. A labeled task 52 clip provides representative Habitat 3.0 visual context; the remaining scene is trace-driven reconstruction.

Episode 75 final state with both task propositions satisfied.
00:00
Act 01 · Task Setup

Move stuffed_toy_0 and toy_vehicle_1 from chair_32 to couch_23.

VIDEO READY · Select an act to seek; playback remains paused

RUN168.15 s

clean seed 0 · λ=0

OBSERVED EVENTstep 820

API mismatch

FORMAL RECOVERYtr=962

Crec=28.4749

FINAL STATEstep 1107

2 / 2 propositions

CASE ATLAS · 05

Stress paths across 10 episodes

180 rollouts support longitudinal comparison across λ and cross-sectional comparison across episodes.

180 / 180completed rollouts
30episode × seed paths
3,115planning events
3task families
PLACEMENT75 · 78 · 98

Fully successful endpoint outcomes still contain distinct recovery costs and paired sensitivities.

18-rollout success
100% for all 3 episodes
β peak
1.905 · Episode 98
SPATIAL RELATION131 · 144 · 174

Outcome variation and support loss expose path-level missingness before family-level aggregation.

Episode 144
66.7% success · 92.2% completion
β peak
3.300 · Episode 144
MIXED / TEMPORAL239 · 258 · 290 · 304

Mixed tasks expose repeated recovery windows, high-stress divergence, and incomplete boundary support.

Episode 239
100% success · 17/18 valid Crec
Episode 304
0% success · 66.7% completion
10-case atlas Episode-level support across 180 rollouts
Each episode contributes 18 rollouts: 3 seeds × 6 stress levels.
FamilyEpisodeSuccessCompletionCrec validRebound pathsβ peakStability pathsGE support
official / strict
Placement75100%100%15/182/30.7902/32/3 · 2/3
Placement78100%100%18/183/31.7463/33/3 · 2/3
Placement98100%100%18/183/31.9053/33/3 · 3/3
Spatial131100%100%0/180/30/30/3 · 0/3
Spatial14466.7%92.2%15/181/33.3001/31/3 · 1/3
Spatial17455.6%77.8%9/180/30/30/3 · 0/3
Mixed239100%100%17/182/31.1512/32/3 · 2/3
Mixed25894.4%95.6%16/181/31.1461/31/3 · 1/3
Mixed29094.4%97.8%17/182/31.4302/32/3 · 2/3
Mixed3040%66.7%0/180/30/30/3 · 0/3