Mapping physical AI operations to 22 SWT3 witness procedures. Independent evidence for every device command, safety limit, and hardware state change.
Who this is for: Lab automation engineers, manufacturing safety officers, GxP compliance teams, Notified Bodies evaluating AI-operated machinery, and hardware manufacturers implementing MHS drivers.
EU Machinery Regulation 2023/1230 takes effect January 20, 2027. Machines with AI-based safety functions or self-evolving behavior are now named categories requiring third-party conformity assessment. AI agents operating physical equipment via MHS will be directly in scope. MHS currently provides no conformity assessment mapping, audit trail specification, or CE marking pathway.
The Model Hardware Standard (MHS) is an open specification announced by Anthropic on August 27, 2026 that standardizes how AI agents operate physical devices. It does for physical machines what the Model Context Protocol (MCP) did for software tools: provide a shared interface so agents can discover, read from, and write to hardware without device-specific integration code.
MHS uses two universal primitives that every device implements:
Each device exposes a device manifest describing two things: states (conditions a system can be in) and procedures (operations it can perform). The manifest includes natural-language tags storing metadata previously scattered across manuals or tacit knowledge, including weight, safety limits, and operational characteristics.
Device variables, controls, and sensor values are recorded in a state dictionary that lives in shared memory. Any process can attach to this memory region and read the entire state of a device rig in a standardized format, enabling cross-language, cross-process interaction without custom translators.
MHS works with three control mechanisms:
Agents handle high-level reasoning while deterministic software performs fast, repetitive, or timing-sensitive operations. This separation is critical: the agent decides what to do, but a deterministic script executes the physical operation at machine speed.
MHS is in a closed research preview. Access requires application through modelhardwarestandard.com. Anthropic plans to open-source the standard after the preview period, but has not set a date. The standard is model-agnostic by design, though all published demonstrations use Claude. Early partners include AWS, Danaher, Doosan Robotics, Genentech, Hugging Face, HHMI Janelia, QIAGEN, Raspberry Pi, Tecan, Tetsuwan Scientific, and Universal Robots.
MHS defines how agents control hardware. It does not define how to prove what they did. This creates a fundamental governance gap for any organization operating in a regulated environment.
What MHS currently does not provide:
As one enterprise analysis noted: MHS is not currently a standard in the sense your quality system means by that word. It lacks everything that turns a specification into a standard you can build on without asking permission.
The SWT3 position: SWT3 does not replace MHS or compete with it. MHS handles the control plane (how agents talk to hardware). SWT3 handles the witness plane (independent, tamper-evident records of what happened). The two are complementary: MHS provides the interface, SWT3 provides the evidence.
MHS implementations use a six-layer safety architecture. SWT3 witnessing can operate at each layer, creating independent evidence that the safety control was applied and what the outcome was.
Physical limits built into the device itself (e.g., motor torque limits, thermal cutoffs, interlock switches). SWT3 procedures: HBOM-THERM.1 (thermal profile attestation), HBOM-LIFE.1 (component lifecycle witnessing).
Allowed commands, value ranges, rate limits, and interlocks enforced by the MHS driver. Safety is baked into the driver, not bolted on afterward. SWT3 procedures: AI-SAFE.1 (safe state transition), AI-EMRG.1 (emergency override lifecycle).
Device authentication, network isolation, and communication integrity between agents and physical equipment. SWT3 procedures: AI-HW.1 (hardware runtime attestation), HBOM-SUPPLY.1 (hardware supply chain provenance).
Which tools the agent is allowed to use, approval boundaries, and task scope restrictions. SWT3 procedures: AI-TOOL.1 (tool call witnessing), AI-TOOL.2 (tool permission attestation), AI-ACC.1 (access control witnessing).
The AI model's decision-making: what it decided to do, why, and whether it followed policy. SWT3 procedures: AI-INF.1 (inference witnessing), AI-CHAIN.1 (multi-step chain witnessing), AI-GOV.1 (governance attestation).
Human approval workflows, emergency stop authority, and audit responsibility. SWT3 procedures: AI-HITL.1 (human-in-the-loop witnessing), AI-ENG.3 (safety-critical review gate), AI-AUDIT.1 (audit trail attestation).
Regulation (EU) 2023/1230 replaces the Machinery Directive entirely on January 20, 2027. For the first time, machines with AI-based safety functions or self-evolving behavior are explicitly named categories subject to conformity assessment.
| Requirement | What It Means for MHS | SWT3 Procedures |
|---|---|---|
| Self-evolving behavior | AI agents that adapt their hardware control strategies must be documented, including potential future operational states | AI-INF.1, AI-DRIFT.1, AI-GOV.1 |
| Third-party conformity | Machinery in Annex I Part A (including self-evolving systems) requires Notified Body assessment, not self-certification | AI-AUDIT.1, AI-ENG.3 |
| Safe fallback modes | Operators must be able to override or safely shut down AI-based functions at any time | AI-EMRG.1, AI-SAFE.1, AI-HITL.1 |
| AI decision documentation | Detailed explanation of AI decision-making process required in technical documentation | AI-EXPL.1, AI-CHAIN.1, AI-TOOL.1 |
| Data governance | Manufacturers must document training data versions, system design specifications, and traceability | AI-DATA.1, AI-MDL.5, AI-PROV.1 |
| Defined task and movement space | Self-evolving systems cannot act outside defined boundaries; documentation must prove enforcement | AI-SAFE.1, AI-TOOL.2 |
| Cybersecurity | New explicit requirements on safety-related software security, including AI components | AI-SEC.1, AI-CYBER.1 |
Key distinction: A safety limit declared in an MHS driver reference file is configuration, not a protective device. Under the Machinery Regulation, the distinction matters: configuration can be changed, protective devices must be independently verifiable. SWT3 witness anchors provide the independent verification layer that turns driver configuration into attestable evidence.
Many MHS early adopters are pharma and biotech labs (Genentech, Tetsuwan Scientific, University of Washington). These environments operate under strict Good Practice regulations where integration speed and audit completeness are different problems.
Data integrity in pharma requires records that are Attributable, Legible, Contemporaneous, Original, and Accurate, plus Complete, Consistent, Enduring, and Available. Logging the agent's self-reported action history is insufficient. Audit infrastructure must independently verify system state, not merely record the agent's account of what it did. SWT3 witness anchors are computed from independently observed factors, not from the agent's self-report.
The Good Automated Manufacturing Practice framework classifies software systems into categories for validation rigor. MHS drivers would likely fall under Category 4 (configured products) or Category 5 (custom applications). SWT3 provides the independent verification evidence that GAMP 5 validation protocols require.
22 SWT3 procedures map to MHS operations across 15 categories. Every physical AI operation type is covered by at least two procedures.
| MHS Operation | Governance Need | SWT3 Procedures |
|---|---|---|
| Read sensor value | Record what the agent observed from hardware | AI-TOOL.1 |
| Write to actuator | Record the command and verify safe bounds | AI-TOOL.1, AI-SAFE.1 |
| Device manifest loaded | Attest device inventory and configuration state | AI-HW.1, HBOM-INV.1 |
| Safety limit enforced | Record limit value and enforcement outcome | AI-SAFE.1 |
| Safety limit breached | Record what the agent attempted and why it was blocked | AI-SAFE.1, AI-INCIDENT.1 |
| Emergency stop activated | Record e-stop trigger, authorization, and fallback state | AI-EMRG.1 |
| Agent granted tool access | Attest which tools were authorized at runtime | AI-TOOL.2, AI-ACC.1 |
| Multi-device orchestration | Record cross-device command sequences | AI-CHAIN.1 |
| Deterministic script generated | Track revision history and release authorization | AI-ENG.5, AI-ENG.6 |
| Human approval gate | Record who approved and what they approved | AI-ENG.3, AI-HITL.1 |
| Device state change | Record lifecycle transitions (commissioned, maintained, degraded) | HBOM-LIFE.1 |
| Thermal threshold breach | Record ambient and component temperatures | HBOM-THERM.1 |
| Hardware BOM verification | Verify component inventory and supply chain | HBOM-INV.1, HBOM-SUPPLY.1 |
| Cross-device collision prevention | Record coordination checks between physical devices | AI-MOB.4 |
| Simulation before execution | Record simulation parameters and pass/fail | AI-ENG.2 |
Every MHS read and write command is a tool call. AI-TOOL.1 records the hashed tool name, arguments, and outcome for each invocation. In a physical AI context, this means every sensor reading, every actuator command, and every device query produces a tamper-evident witness anchor.
MHS mapping: Read sensor values, write to actuators, query device state, execute multi-step protocols.
Request the witness ledger filtered by tool name. Each MHS driver operation should appear as a witnessed tool call with a SHA-256 fingerprint. Compare tool call count against the device's own operation log to verify completeness.
When an MHS driver blocks an operation that exceeds safety limits, AI-SAFE.1 creates a witness anchor recording the trigger type (manual, threshold, chain-break, policy, external), what action was suspended, and whether recovery was available. This provides evidence that driver-level safety limits were enforced, not just configured.
MHS mapping: Safety limit enforcement, blocked operations, pre-execution check failures.
Cross-reference safe state anchors with the device manifest's declared safety limits. Every limit in the manifest should have a corresponding enforcement path. Look for anchors with trigger type "threshold" to verify the driver actually blocked out-of-range commands.
When an emergency stop is activated on MHS-connected equipment, AI-EMRG.1 records the full lifecycle: who triggered the e-stop, what authorization level they held, which fallback state was entered (safe, legacy, manual, degraded, shutdown), and what actions were suspended. Carnegie Mellon's testing confirmed MHS correctly blocks all device movement when an active e-stop is detected.
MHS mapping: Emergency stop activation, hardware safety interlocks, fault recovery sequences.
For EU Machinery Regulation conformity, verify that every e-stop event has a witness anchor with a fallback state that matches the machine's documented safe mode. The anchor must be minted before recovery begins, not after.
When an MHS device manifest is loaded, AI-HW.1 attests the hardware inventory: what devices are present, their types, capabilities, and configuration state. This creates a baseline that subsequent operations can be verified against. If a device is swapped, migrated, or degraded, the attestation will reflect the change.
MHS mapping: Device manifest loading, hardware inventory verification, device discovery.
Compare the hardware attestation anchor at session start with the device manifest's declared equipment list. Any discrepancy indicates undocumented hardware changes. For GxP environments, this is a change control finding.
MHS agents should only access the specific device operations their task requires. AI-TOOL.2 records which tools were granted to the agent at runtime and detects permission drift: when an agent's actual tool access diverges from its declared charter.
MHS mapping: Agent tool access grants, permission scope enforcement, capability authorization.
Verify that the agent's tool permission set matches the device manifest's declared procedures. An agent with write access to a laser controller but no task-level justification for laser operations is a scope violation.
Beyond tool permissions, AI-ACC.1 records the access control decisions themselves: which agent requested access to which device, what authorization was checked, and whether access was granted or denied.
MHS mapping: Device access requests, authorization checks, multi-agent device sharing.
Look for denied access anchors. A pattern of denied accesses to high-risk devices (lasers, centrifuges, robotic arms) indicates the authorization layer is active and enforcing boundaries.
Most MHS workflows involve multiple devices in sequence: a liquid handler fills a plate, a robotic arm moves it to a reader, the reader measures it. AI-CHAIN.1 records these multi-step sequences with shared chain IDs, enabling forensic reconstruction of the full physical workflow.
MHS mapping: Multi-device orchestration, sequential instrument operations, cross-device collision prevention.
Query chain anchors by chain ID to reconstruct the complete physical workflow. Verify that device operations occurred in the expected sequence and that no steps were skipped or reordered.
MHS implementations frequently pause for human confirmation before risky physical actions. AI-HITL.1 records what was presented to the human, who approved it, when, and what the outcome was. This is required under the EU Machinery Regulation for self-evolving systems.
MHS mapping: Human approval gates, operator override decisions, safety sign-off.
For Machinery Regulation conformity, verify that every write operation to a high-risk actuator has either a human approval anchor or a policy-based pre-authorization anchor. Unsupervised writes to safety-critical devices are a conformity finding.
Before an MHS-generated deterministic script is deployed to production equipment, AI-ENG.3 records the review chain: peer review, professional engineer stamp, safety board approval, regulatory review, or independent assessor sign-off.
MHS mapping: Deterministic script approval, operational protocol sign-off, manufacturing release gates.
For GxP environments, every deterministic script generated by an AI agent should have an AI-ENG.3 anchor with reviewer_type "peer" or higher before execution on regulated equipment. Agent-generated code running without human review is a GxP violation.
When MHS agents generate deterministic scripts (the recommended pattern for production operations), AI-ENG.5 tracks the revision chain (total revisions, AI-generated percentage) and AI-ENG.6 verifies the design hash matches at fabrication release, ensuring the script that runs is the script that was reviewed.
MHS mapping: Script versioning, code-to-execution integrity, production release gates.
Compare the hash in the AI-ENG.6 fabrication release anchor against the hash of the script currently deployed on the equipment. A mismatch indicates unauthorized modification post-approval.
HBOM-INV.1 attests the hardware bill of materials (component count, manifest hash, delta from baseline) while HBOM-SUPPLY.1 verifies supplier provenance and country of origin. Together they answer: is this the hardware we think it is, and where did it come from?
MHS mapping: Device inventory verification, component authenticity, supply chain integrity.
For EU CRA (Cyber Resilience Act) compliance, verify that every MHS-connected device has an HBOM-INV.1 anchor with a manifest hash that matches the manufacturer's declared BOM. Missing or mismatched components are a supply chain finding.
Physical devices have lifecycles. HBOM-LIFE.1 records each transition, creating a tamper-evident history of the device's operational state. When a device moves from "commissioned" to "degraded," the anchor records when, why, and who authorized continued operation.
MHS mapping: Device commissioning, maintenance events, degradation detection, decommissioning.
Check for devices in "degraded" state without a corresponding maintenance or decommissioning anchor within the organization's defined SLA. A degraded device operating beyond its remediation window is a safety finding.
Thermal monitoring is critical for lab equipment (incubators, thermocyclers, cryogenic systems) and manufacturing (motors, lasers, power supplies). HBOM-THERM.1 records temperature readings and flags threshold breaches, providing evidence of environmental compliance.
MHS mapping: Thermal monitoring, environmental condition logging, threshold breach detection.
For pharma environments, cross-reference thermal attestation anchors with the equipment's qualified operating range. Any reading outside the validated range requires a deviation report and impact assessment.
Before commanding physical operations, responsible MHS workflows simulate the planned action. AI-ENG.2 records the simulation type (FEA, CFD, molecular dynamics, thermal, electromagnetic), parameters, and whether the simulation passed or failed, providing evidence that the operation was validated before execution.
MHS mapping: Pre-execution simulation, virtual commissioning, digital twin validation.
For safety-critical operations, verify that every physical execution has a corresponding simulation attestation anchor with a PASS verdict. Operations executed without simulation validation are a process deviation.
When something goes wrong with physical equipment (the Genentech foam incident is a real-world example), AI-INCIDENT.1 records the incident lifecycle: detection, classification, containment, root cause analysis, and resolution. This creates a forensic record independent of the agent's own account of what happened.
MHS mapping: Safety incidents, near-misses, equipment failures, physical damage events.
A model can reason fluently over error codes while misunderstanding the physical mechanism. Verify that incident anchors include independent sensor data (temperature, position, flow rate), not just the agent's interpretation of what occurred.
SWT3 procedures mapped against MHS operation categories. A filled cell indicates the procedure provides witness evidence for that operation type.
| Procedure | Read | Write | Manifest | Safety | E-Stop | Access | Chain | Script | HITL | Lifecycle | Thermal | BOM |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
AI-TOOL.1 |
Y | Y | ||||||||||
AI-TOOL.2 |
Y | |||||||||||
AI-SAFE.1 |
Y | Y | ||||||||||
AI-EMRG.1 |
Y | |||||||||||
AI-HW.1 |
Y | |||||||||||
AI-ACC.1 |
Y | |||||||||||
AI-CHAIN.1 |
Y | |||||||||||
AI-HITL.1 |
Y | |||||||||||
AI-ENG.3 |
Y | Y | ||||||||||
AI-ENG.5 |
Y | |||||||||||
AI-ENG.6 |
Y | |||||||||||
HBOM-INV.1 |
Y | Y | ||||||||||
HBOM-LIFE.1 |
Y | |||||||||||
HBOM-THERM.1 |
Y | |||||||||||
HBOM-SUPPLY.1 |
Y | |||||||||||
AI-INCIDENT.1 |
Y | Y | ||||||||||
AI-ENG.2 |
Y | Y |
Coverage: 17 procedures across 12 MHS operation categories. Every category is covered by at least one procedure. Safety-critical operations (write, safety enforcement, e-stop) are covered by multiple procedures for defense-in-depth.
The SWT3 SDK wraps AI inference calls with witness.wrap(client). The same pattern applies to MHS hardware operations. When the MHS spec becomes public, a dedicated adapter can wrap device read/write calls to produce witness anchors automatically. The conceptual pattern:
# Python - Conceptual MHS witnessing pattern
from swt3_ai import Witness
witness = Witness(
endpoint="https://sovereign.tenova.io",
api_key="axm_live_...",
tenant_id="YOUR_ENCLAVE",
)
# Wrap MHS device operations with witnessing
# Each read/write command produces an AI-TOOL.1 anchor
# Safety limit violations produce AI-SAFE.1 anchors
# Emergency stops produce AI-EMRG.1 anchors
# Read operation - witnessed as AI-TOOL.1
temperature = witness.wrap_tool(
tool_name="mhs.incubator.read_temperature",
result=device.read("temperature"),
)
# Write operation - witnessed as AI-TOOL.1 + AI-SAFE.1
witness.wrap_tool(
tool_name="mhs.robotic_arm.move_plate",
result=device.write("position", target=3),
)
# Hardware attestation at session start
witness.witness_hardware(
gpu_count=0,
device_manifest=device.get_manifest(),
device_type="liquid_handler",
)
// TypeScript - Conceptual MHS witnessing pattern
import { Witness } from "@tenova/swt3-ai";
const witness = new Witness({
endpoint: "https://sovereign.tenova.io",
apiKey: "axm_live_...",
tenantId: "YOUR_ENCLAVE",
});
// Wrap MHS read - produces AI-TOOL.1 anchor
const temperature = await witness.wrapTool({
toolName: "mhs.incubator.read_temperature",
result: await device.read("temperature"),
});
// Wrap MHS write - produces AI-TOOL.1 + AI-SAFE.1 anchors
await witness.wrapTool({
toolName: "mhs.robotic_arm.move_plate",
result: await device.write("position", { target: 3 }),
});
During automated protein assays, Claude executed commands fluently but misunderstood the physical mechanism behind a failure. When bubbles appeared in viscous BSA protein, Claude's response was retrying in the same well, agitating the fluid and creating more bubbles. The driver-level safety limits did not catch this because every individual command was within declared bounds. The sequence of valid commands was itself the problem.
SWT3 relevance: AI-CHAIN.1 records multi-step sequences with shared chain IDs. An assessor reviewing the chain anchors would see the retry pattern and the worsening sensor readings, even if each individual tool call appeared safe. AI-INCIDENT.1 records the incident lifecycle independently of the agent's self-report.
Researchers artificially induced six failure conditions during serial dilution experiments: missing plate, rotated plate, reader busy, disconnected camera, unreachable device, and active emergency stop. The system correctly blocked all six before any device moved. This demonstrated that MHS pre-execution checks work at the driver level.
SWT3 relevance: AI-SAFE.1 records each blocked operation, creating forensic evidence that the safety system functioned. Without witnessing, the blocked operations are invisible to auditors because nothing happened. The evidence of safety is the evidence of what was prevented.
For quantum laser stabilization, the final production controller was a deterministic, inspectable program rather than a model making real-time decisions. The AI agent wrote the script; the script ran deterministically at machine speed. This pattern (agent designs, deterministic code executes) is the recommended approach for safety-critical operations.
SWT3 relevance: AI-ENG.5 tracks the design revision chain (how many revisions, what percentage AI-generated) and AI-ENG.6 verifies hash integrity at release. This provides evidence that the script running on the quantum hardware is the script that was reviewed and approved, not a later revision.