A safety filter reins in residual reinforcement learning for simulated spacecraft formations
The authors combine a nominal controller, a structured residual policy, and a recursive online safety filter for coupled translation-and-attitude control. The available abstract reports better tracking and training convergence than pure policy learning and single-step shielding, with near-zero simulated safety violations—but supplies no numerical results or evidence of flight or hardware testing.
Boundary: Publisher Cloudflare controls blocked direct inspection of the landing page and PDF, leaving equations, proofs, tables, figures, numerical results, and experimental details unavailable. The number of spacecraft, orbit regime, simulator, scenarios, episodes, seeds, uncertainty distributions, comparator implementations, filter-intervention rates, runtimes, ablations, and out-of-distribution tests could not be verified. The abstract reports simulations only—no hardware-in-the-loop, air-bearing, robotic, or on-orbit experiment. Near-zero violations are not violation-free performance, and moderate computation cost does not establish compatibility with an onboard processor or control cycle. No code, data, independent replication, or consolidated artifact was located.