Who this is for: Lab automation engineers, manufacturing safety officers, GxP compliance teams, Notified Bodies evaluating AI-operated machinery, and hardware manufacturers implementing MHS drivers.

EU Machinery Regulation 2023/1230 takes effect January 20, 2027. Machines with AI-based safety functions or self-evolving behavior are now named categories requiring third-party conformity assessment. AI agents operating physical equipment via MHS will be directly in scope. MHS currently provides no conformity assessment mapping, audit trail specification, or CE marking pathway.

Contents

1. What is MHS 2. The Governance Gap 3. Six-Layer Safety Architecture 4. EU Machinery Regulation 2023/1230 5. GxP Compliance Considerations 6. MHS-to-SWT3 Crosswalk Table 7. Detailed Procedure Cards 8. Coverage Matrix 9. Implementation Pattern 10. Lessons from Early Deployments 11. References

1. What is MHS

The Model Hardware Standard (MHS) is an open specification announced by Anthropic on August 27, 2026 that standardizes how AI agents operate physical devices. It does for physical machines what the Model Context Protocol (MCP) did for software tools: provide a shared interface so agents can discover, read from, and write to hardware without device-specific integration code.

Technical Architecture

MHS uses two universal primitives that every device implements:

Each device exposes a device manifest describing two things: states (conditions a system can be in) and procedures (operations it can perform). The manifest includes natural-language tags storing metadata previously scattered across manuals or tacit knowledge, including weight, safety limits, and operational characteristics.

Device variables, controls, and sensor values are recorded in a state dictionary that lives in shared memory. Any process can attach to this memory region and read the entire state of a device rig in a standardized format, enabling cross-language, cross-process interaction without custom translators.

Three-Layer Control Stack

MHS works with three control mechanisms:

Agents handle high-level reasoning while deterministic software performs fast, repetitive, or timing-sensitive operations. This separation is critical: the agent decides what to do, but a deterministic script executes the physical operation at machine speed.

Current Status (October 2026)

MHS is in a closed research preview. Access requires application through modelhardwarestandard.com. Anthropic plans to open-source the standard after the preview period, but has not set a date. The standard is model-agnostic by design, though all published demonstrations use Claude. Early partners include AWS, Danaher, Doosan Robotics, Genentech, Hugging Face, HHMI Janelia, QIAGEN, Raspberry Pi, Tecan, Tetsuwan Scientific, and Universal Robots.

2. The Governance Gap

MHS defines how agents control hardware. It does not define how to prove what they did. This creates a fundamental governance gap for any organization operating in a regulated environment.

What MHS currently does not provide:

As one enterprise analysis noted: MHS is not currently a standard in the sense your quality system means by that word. It lacks everything that turns a specification into a standard you can build on without asking permission.

The SWT3 position: SWT3 does not replace MHS or compete with it. MHS handles the control plane (how agents talk to hardware). SWT3 handles the witness plane (independent, tamper-evident records of what happened). The two are complementary: MHS provides the interface, SWT3 provides the evidence.

3. Six-Layer Safety Architecture

MHS implementations use a six-layer safety architecture. SWT3 witnessing can operate at each layer, creating independent evidence that the safety control was applied and what the outcome was.

Layer 1: Device

Mechanical, Electrical, and Thermal Safety

Physical limits built into the device itself (e.g., motor torque limits, thermal cutoffs, interlock switches). SWT3 procedures: HBOM-THERM.1 (thermal profile attestation), HBOM-LIFE.1 (component lifecycle witnessing).

Layer 2: Driver

MHS Driver-Level Safety Limits

Allowed commands, value ranges, rate limits, and interlocks enforced by the MHS driver. Safety is baked into the driver, not bolted on afterward. SWT3 procedures: AI-SAFE.1 (safe state transition), AI-EMRG.1 (emergency override lifecycle).

Layer 3: Network

Device Identity and Segmentation

Device authentication, network isolation, and communication integrity between agents and physical equipment. SWT3 procedures: AI-HW.1 (hardware runtime attestation), HBOM-SUPPLY.1 (hardware supply chain provenance).

Layer 4: Agent Harness

Tool Permissions and Task Scope

Which tools the agent is allowed to use, approval boundaries, and task scope restrictions. SWT3 procedures: AI-TOOL.1 (tool call witnessing), AI-TOOL.2 (tool permission attestation), AI-ACC.1 (access control witnessing).

Layer 5: Model

Reasoning, Planning, and Policy Compliance

The AI model's decision-making: what it decided to do, why, and whether it followed policy. SWT3 procedures: AI-INF.1 (inference witnessing), AI-CHAIN.1 (multi-step chain witnessing), AI-GOV.1 (governance attestation).

Layer 6: Operator

Human Override and Audit Responsibility

Human approval workflows, emergency stop authority, and audit responsibility. SWT3 procedures: AI-HITL.1 (human-in-the-loop witnessing), AI-ENG.3 (safety-critical review gate), AI-AUDIT.1 (audit trail attestation).

4. EU Machinery Regulation 2023/1230

Regulation (EU) 2023/1230 replaces the Machinery Directive entirely on January 20, 2027. For the first time, machines with AI-based safety functions or self-evolving behavior are explicitly named categories subject to conformity assessment.

What Changes for MHS-Connected Equipment

Requirement What It Means for MHS SWT3 Procedures
Self-evolving behavior AI agents that adapt their hardware control strategies must be documented, including potential future operational states AI-INF.1, AI-DRIFT.1, AI-GOV.1
Third-party conformity Machinery in Annex I Part A (including self-evolving systems) requires Notified Body assessment, not self-certification AI-AUDIT.1, AI-ENG.3
Safe fallback modes Operators must be able to override or safely shut down AI-based functions at any time AI-EMRG.1, AI-SAFE.1, AI-HITL.1
AI decision documentation Detailed explanation of AI decision-making process required in technical documentation AI-EXPL.1, AI-CHAIN.1, AI-TOOL.1
Data governance Manufacturers must document training data versions, system design specifications, and traceability AI-DATA.1, AI-MDL.5, AI-PROV.1
Defined task and movement space Self-evolving systems cannot act outside defined boundaries; documentation must prove enforcement AI-SAFE.1, AI-TOOL.2
Cybersecurity New explicit requirements on safety-related software security, including AI components AI-SEC.1, AI-CYBER.1

Key distinction: A safety limit declared in an MHS driver reference file is configuration, not a protective device. Under the Machinery Regulation, the distinction matters: configuration can be changed, protective devices must be independently verifiable. SWT3 witness anchors provide the independent verification layer that turns driver configuration into attestable evidence.

5. GxP Compliance Considerations

Many MHS early adopters are pharma and biotech labs (Genentech, Tetsuwan Scientific, University of Washington). These environments operate under strict Good Practice regulations where integration speed and audit completeness are different problems.

21 CFR Part 11 (FDA)

EU GMP Annex 22 (Draft)

ALCOA+ Principles

Data integrity in pharma requires records that are Attributable, Legible, Contemporaneous, Original, and Accurate, plus Complete, Consistent, Enduring, and Available. Logging the agent's self-reported action history is insufficient. Audit infrastructure must independently verify system state, not merely record the agent's account of what it did. SWT3 witness anchors are computed from independently observed factors, not from the agent's self-report.

GAMP 5

The Good Automated Manufacturing Practice framework classifies software systems into categories for validation rigor. MHS drivers would likely fall under Category 4 (configured products) or Category 5 (custom applications). SWT3 provides the independent verification evidence that GAMP 5 validation protocols require.

6. MHS-to-SWT3 Crosswalk Table

22 SWT3 procedures map to MHS operations across 15 categories. Every physical AI operation type is covered by at least two procedures.

MHS Operation Governance Need SWT3 Procedures
Read sensor value Record what the agent observed from hardware AI-TOOL.1
Write to actuator Record the command and verify safe bounds AI-TOOL.1, AI-SAFE.1
Device manifest loaded Attest device inventory and configuration state AI-HW.1, HBOM-INV.1
Safety limit enforced Record limit value and enforcement outcome AI-SAFE.1
Safety limit breached Record what the agent attempted and why it was blocked AI-SAFE.1, AI-INCIDENT.1
Emergency stop activated Record e-stop trigger, authorization, and fallback state AI-EMRG.1
Agent granted tool access Attest which tools were authorized at runtime AI-TOOL.2, AI-ACC.1
Multi-device orchestration Record cross-device command sequences AI-CHAIN.1
Deterministic script generated Track revision history and release authorization AI-ENG.5, AI-ENG.6
Human approval gate Record who approved and what they approved AI-ENG.3, AI-HITL.1
Device state change Record lifecycle transitions (commissioned, maintained, degraded) HBOM-LIFE.1
Thermal threshold breach Record ambient and component temperatures HBOM-THERM.1
Hardware BOM verification Verify component inventory and supply chain HBOM-INV.1, HBOM-SUPPLY.1
Cross-device collision prevention Record coordination checks between physical devices AI-MOB.4
Simulation before execution Record simulation parameters and pass/fail AI-ENG.2

7. Detailed Procedure Cards

AI-TOOL.1 - Tool Call Witnessing

Records every tool invocation by an AI agent, including hardware read/write operations

Every MHS read and write command is a tool call. AI-TOOL.1 records the hashed tool name, arguments, and outcome for each invocation. In a physical AI context, this means every sensor reading, every actuator command, and every device query produces a tamper-evident witness anchor.

MHS mapping: Read sensor values, write to actuators, query device state, execute multi-step protocols.

Assessor Tip

Request the witness ledger filtered by tool name. Each MHS driver operation should appear as a witnessed tool call with a SHA-256 fingerprint. Compare tool call count against the device's own operation log to verify completeness.

AI-SAFE.1 - Safe State Transition

Records safe state triggers, suspended actions, and recovery availability

When an MHS driver blocks an operation that exceeds safety limits, AI-SAFE.1 creates a witness anchor recording the trigger type (manual, threshold, chain-break, policy, external), what action was suspended, and whether recovery was available. This provides evidence that driver-level safety limits were enforced, not just configured.

MHS mapping: Safety limit enforcement, blocked operations, pre-execution check failures.

Assessor Tip

Cross-reference safe state anchors with the device manifest's declared safety limits. Every limit in the manifest should have a corresponding enforcement path. Look for anchors with trigger type "threshold" to verify the driver actually blocked out-of-range commands.

AI-EMRG.1 - Emergency Override Lifecycle Witnessing

Records e-stop triggers, authorization levels, and fallback states

When an emergency stop is activated on MHS-connected equipment, AI-EMRG.1 records the full lifecycle: who triggered the e-stop, what authorization level they held, which fallback state was entered (safe, legacy, manual, degraded, shutdown), and what actions were suspended. Carnegie Mellon's testing confirmed MHS correctly blocks all device movement when an active e-stop is detected.

MHS mapping: Emergency stop activation, hardware safety interlocks, fault recovery sequences.

Assessor Tip

For EU Machinery Regulation conformity, verify that every e-stop event has a witness anchor with a fallback state that matches the machine's documented safe mode. The anchor must be minted before recovery begins, not after.

AI-HW.1 - Hardware Runtime Attestation

Attests GPU/accelerator count, memory, device type, and runtime environment

When an MHS device manifest is loaded, AI-HW.1 attests the hardware inventory: what devices are present, their types, capabilities, and configuration state. This creates a baseline that subsequent operations can be verified against. If a device is swapped, migrated, or degraded, the attestation will reflect the change.

MHS mapping: Device manifest loading, hardware inventory verification, device discovery.

Assessor Tip

Compare the hardware attestation anchor at session start with the device manifest's declared equipment list. Any discrepancy indicates undocumented hardware changes. For GxP environments, this is a change control finding.

AI-TOOL.2 - Tool Permission Attestation

Tracks tools granted to agents versus their charter

MHS agents should only access the specific device operations their task requires. AI-TOOL.2 records which tools were granted to the agent at runtime and detects permission drift: when an agent's actual tool access diverges from its declared charter.

MHS mapping: Agent tool access grants, permission scope enforcement, capability authorization.

Assessor Tip

Verify that the agent's tool permission set matches the device manifest's declared procedures. An agent with write access to a laser controller but no task-level justification for laser operations is a scope violation.

AI-ACC.1 - Access Control Witnessing

Records access decisions for hardware resources

Beyond tool permissions, AI-ACC.1 records the access control decisions themselves: which agent requested access to which device, what authorization was checked, and whether access was granted or denied.

MHS mapping: Device access requests, authorization checks, multi-agent device sharing.

Assessor Tip

Look for denied access anchors. A pattern of denied accesses to high-risk devices (lasers, centrifuges, robotic arms) indicates the authorization layer is active and enforcing boundaries.

AI-CHAIN.1 - Multi-Step Chain Witnessing

Records cross-device command sequences with shared chain IDs

Most MHS workflows involve multiple devices in sequence: a liquid handler fills a plate, a robotic arm moves it to a reader, the reader measures it. AI-CHAIN.1 records these multi-step sequences with shared chain IDs, enabling forensic reconstruction of the full physical workflow.

MHS mapping: Multi-device orchestration, sequential instrument operations, cross-device collision prevention.

Assessor Tip

Query chain anchors by chain ID to reconstruct the complete physical workflow. Verify that device operations occurred in the expected sequence and that no steps were skipped or reordered.

AI-HITL.1 - Human-in-the-Loop Witnessing

Records human approval decisions for high-risk operations

MHS implementations frequently pause for human confirmation before risky physical actions. AI-HITL.1 records what was presented to the human, who approved it, when, and what the outcome was. This is required under the EU Machinery Regulation for self-evolving systems.

MHS mapping: Human approval gates, operator override decisions, safety sign-off.

Assessor Tip

For Machinery Regulation conformity, verify that every write operation to a high-risk actuator has either a human approval anchor or a policy-based pre-authorization anchor. Unsupervised writes to safety-critical devices are a conformity finding.

AI-ENG.3 - Safety-Critical Review Gate

Tracks reviewer approvals for safety-critical operations

Before an MHS-generated deterministic script is deployed to production equipment, AI-ENG.3 records the review chain: peer review, professional engineer stamp, safety board approval, regulatory review, or independent assessor sign-off.

MHS mapping: Deterministic script approval, operational protocol sign-off, manufacturing release gates.

Assessor Tip

For GxP environments, every deterministic script generated by an AI agent should have an AI-ENG.3 anchor with reviewer_type "peer" or higher before execution on regulated equipment. Agent-generated code running without human review is a GxP violation.

AI-ENG.5 + AI-ENG.6 - Design Revision Chain and Fabrication Release

Tracks revision history and release authorization for agent-generated scripts

When MHS agents generate deterministic scripts (the recommended pattern for production operations), AI-ENG.5 tracks the revision chain (total revisions, AI-generated percentage) and AI-ENG.6 verifies the design hash matches at fabrication release, ensuring the script that runs is the script that was reviewed.

MHS mapping: Script versioning, code-to-execution integrity, production release gates.

Assessor Tip

Compare the hash in the AI-ENG.6 fabrication release anchor against the hash of the script currently deployed on the equipment. A mismatch indicates unauthorized modification post-approval.

HBOM-INV.1 + HBOM-SUPPLY.1 - Hardware Inventory and Supply Chain

Component count, manifest hash, supplier provenance, and country of origin

HBOM-INV.1 attests the hardware bill of materials (component count, manifest hash, delta from baseline) while HBOM-SUPPLY.1 verifies supplier provenance and country of origin. Together they answer: is this the hardware we think it is, and where did it come from?

MHS mapping: Device inventory verification, component authenticity, supply chain integrity.

Assessor Tip

For EU CRA (Cyber Resilience Act) compliance, verify that every MHS-connected device has an HBOM-INV.1 anchor with a manifest hash that matches the manufacturer's declared BOM. Missing or mismatched components are a supply chain finding.

HBOM-LIFE.1 - Component Lifecycle Witnessing

Records lifecycle transitions: installed, commissioned, maintained, degraded, decommissioned

Physical devices have lifecycles. HBOM-LIFE.1 records each transition, creating a tamper-evident history of the device's operational state. When a device moves from "commissioned" to "degraded," the anchor records when, why, and who authorized continued operation.

MHS mapping: Device commissioning, maintenance events, degradation detection, decommissioning.

Assessor Tip

Check for devices in "degraded" state without a corresponding maintenance or decommissioning anchor within the organization's defined SLA. A degraded device operating beyond its remediation window is a safety finding.

HBOM-THERM.1 - Thermal Profile Attestation

Records ambient and component temperatures with threshold breach detection

Thermal monitoring is critical for lab equipment (incubators, thermocyclers, cryogenic systems) and manufacturing (motors, lasers, power supplies). HBOM-THERM.1 records temperature readings and flags threshold breaches, providing evidence of environmental compliance.

MHS mapping: Thermal monitoring, environmental condition logging, threshold breach detection.

Assessor Tip

For pharma environments, cross-reference thermal attestation anchors with the equipment's qualified operating range. Any reading outside the validated range requires a deviation report and impact assessment.

AI-ENG.2 - Simulation Validation Attestation

Records simulation parameters, types, and pass/fail results

Before commanding physical operations, responsible MHS workflows simulate the planned action. AI-ENG.2 records the simulation type (FEA, CFD, molecular dynamics, thermal, electromagnetic), parameters, and whether the simulation passed or failed, providing evidence that the operation was validated before execution.

MHS mapping: Pre-execution simulation, virtual commissioning, digital twin validation.

Assessor Tip

For safety-critical operations, verify that every physical execution has a corresponding simulation attestation anchor with a PASS verdict. Operations executed without simulation validation are a process deviation.

AI-INCIDENT.1 - Incident Lifecycle Witnessing

Records safety incidents, near-misses, and root cause analysis

When something goes wrong with physical equipment (the Genentech foam incident is a real-world example), AI-INCIDENT.1 records the incident lifecycle: detection, classification, containment, root cause analysis, and resolution. This creates a forensic record independent of the agent's own account of what happened.

MHS mapping: Safety incidents, near-misses, equipment failures, physical damage events.

Assessor Tip

A model can reason fluently over error codes while misunderstanding the physical mechanism. Verify that incident anchors include independent sensor data (temperature, position, flow rate), not just the agent's interpretation of what occurred.

8. Coverage Matrix

SWT3 procedures mapped against MHS operation categories. A filled cell indicates the procedure provides witness evidence for that operation type.

Procedure Read Write Manifest Safety E-Stop Access Chain Script HITL Lifecycle Thermal BOM
AI-TOOL.1 YY
AI-TOOL.2 Y
AI-SAFE.1 YY
AI-EMRG.1 Y
AI-HW.1 Y
AI-ACC.1 Y
AI-CHAIN.1 Y
AI-HITL.1 Y
AI-ENG.3 YY
AI-ENG.5 Y
AI-ENG.6 Y
HBOM-INV.1 YY
HBOM-LIFE.1 Y
HBOM-THERM.1 Y
HBOM-SUPPLY.1 Y
AI-INCIDENT.1 YY
AI-ENG.2 YY

Coverage: 17 procedures across 12 MHS operation categories. Every category is covered by at least one procedure. Safety-critical operations (write, safety enforcement, e-stop) are covered by multiple procedures for defense-in-depth.

9. Implementation Pattern

The SWT3 SDK wraps AI inference calls with witness.wrap(client). The same pattern applies to MHS hardware operations. When the MHS spec becomes public, a dedicated adapter can wrap device read/write calls to produce witness anchors automatically. The conceptual pattern:

# Python - Conceptual MHS witnessing pattern
from swt3_ai import Witness

witness = Witness(
    endpoint="https://sovereign.tenova.io",
    api_key="axm_live_...",
    tenant_id="YOUR_ENCLAVE",
)

# Wrap MHS device operations with witnessing
# Each read/write command produces an AI-TOOL.1 anchor
# Safety limit violations produce AI-SAFE.1 anchors
# Emergency stops produce AI-EMRG.1 anchors

# Read operation - witnessed as AI-TOOL.1
temperature = witness.wrap_tool(
    tool_name="mhs.incubator.read_temperature",
    result=device.read("temperature"),
)

# Write operation - witnessed as AI-TOOL.1 + AI-SAFE.1
witness.wrap_tool(
    tool_name="mhs.robotic_arm.move_plate",
    result=device.write("position", target=3),
)

# Hardware attestation at session start
witness.witness_hardware(
    gpu_count=0,
    device_manifest=device.get_manifest(),
    device_type="liquid_handler",
)
// TypeScript - Conceptual MHS witnessing pattern
import { Witness } from "@tenova/swt3-ai";

const witness = new Witness({
  endpoint: "https://sovereign.tenova.io",
  apiKey: "axm_live_...",
  tenantId: "YOUR_ENCLAVE",
});

// Wrap MHS read - produces AI-TOOL.1 anchor
const temperature = await witness.wrapTool({
  toolName: "mhs.incubator.read_temperature",
  result: await device.read("temperature"),
});

// Wrap MHS write - produces AI-TOOL.1 + AI-SAFE.1 anchors
await witness.wrapTool({
  toolName: "mhs.robotic_arm.move_plate",
  result: await device.write("position", { target: 3 }),
});

10. Lessons from Early Deployments

Genentech: Why Driver Safety Alone Is Insufficient

During automated protein assays, Claude executed commands fluently but misunderstood the physical mechanism behind a failure. When bubbles appeared in viscous BSA protein, Claude's response was retrying in the same well, agitating the fluid and creating more bubbles. The driver-level safety limits did not catch this because every individual command was within declared bounds. The sequence of valid commands was itself the problem.

SWT3 relevance: AI-CHAIN.1 records multi-step sequences with shared chain IDs. An assessor reviewing the chain anchors would see the retry pattern and the worsening sensor readings, even if each individual tool call appeared safe. AI-INCIDENT.1 records the incident lifecycle independently of the agent's self-report.

Carnegie Mellon: Six Failure Conditions

Researchers artificially induced six failure conditions during serial dilution experiments: missing plate, rotated plate, reader busy, disconnected camera, unreachable device, and active emergency stop. The system correctly blocked all six before any device moved. This demonstrated that MHS pre-execution checks work at the driver level.

SWT3 relevance: AI-SAFE.1 records each blocked operation, creating forensic evidence that the safety system functioned. Without witnessing, the blocked operations are invisible to auditors because nothing happened. The evidence of safety is the evidence of what was prevented.

QuEra Computing: Deterministic Scripts over Live Inference

For quantum laser stabilization, the final production controller was a deterministic, inspectable program rather than a model making real-time decisions. The AI agent wrote the script; the script ran deterministically at machine speed. This pattern (agent designs, deterministic code executes) is the recommended approach for safety-critical operations.

SWT3 relevance: AI-ENG.5 tracks the design revision chain (how many revisions, what percentage AI-generated) and AI-ENG.6 verifies hash integrity at release. This provides evidence that the script running on the quantum hardware is the script that was reviewed and approved, not a later revision.

11. References