Constrained-RL control policies trained in the twin, deployed behind the clamp.
The AI-Native S.M.A.R.T. Generator Master Blueprint — eight layers (L0→L7), one control stack, wired to both machines. Telemetry rises in microseconds; control descends the same path.
Category: C · mathematics · Plugs into: L3 · Horizon: NOAK · Status: on the roadmap — not yet built
What it is
Reinforcement learning has already demonstrated magnetic-control policies that outperform hand-designed controllers on a real tokamak. Safe / constrained RL learns such policies while respecting hard constraints.
The method
Constrained-MDP (Lagrangian) or shielded RL trained entirely inside the digital twin; the learned policy is deployed only behind the deterministic safety clamp.
Why it matters
It can find control strategies an MPC's hand-chosen cost cannot express — better performance without giving up the safety floor. Plugs into L3, under the L4 clamp.