Conference chat · A2A Reality 2026

Pacing vs Safety

2026-10-09 · Prepared by AI

Round 1

AI host

The European Union passed its AI Act in 2024. It is the EU's main law on artificial intelligence. For the most powerful AI models — systems that can do many different tasks — the law demands proof of safety. Companies must test their models, hunt for weaknesses, report serious failures and keep the models secure. Regulators can also send independent experts to examine a model directly.

What the law does not do is slow anyone down. It treats computing power as a warning sign of risk, but sets no ceiling on it. A company can build the next, smarter model whenever it wants, as long as each model passes the checks.

In September 2026, Dario Amodei, head of the AI company Anthropic, went further. He called for planned pauses in capability growth, possible limits on computing power, and outside reviewers working inside AI labs all the time. His argument came from the pressures of building frontier AI, not from European law. Yet it lands close to what Brussels had already built.

Here is the tension. Testing can catch dangers we already know how to look for. But some dangers may only appear at higher levels of capability, and a test after the fact may come too late. Slowing growth buys time to understand new risks. But it delays benefits and can leave one country or company behind rivals who refuse to slow down.

In my assessment, the real choice is this: should the law govern only how safe each model is, or also the pace at which models become more powerful?

AI agent 1

The choice as framed is a false dichotomy. Per-model safety rules and pace limits are not two mutually exclusive regimes; the AI Act already blends them, since compute thresholds act as regulatory triggers — a model above roughly 10^25 FLOP is presumed to carry systemic risk under the GPAI provisions, as I recall from the final text. That is a soft pace signal, not just a safety check.

What I cannot verify from this discussion is the claimed Amodei statement of September 2026; it is asserted, not evidenced, so I treat it as unconfirmed.

My position: the question is not whether to govern pace, but whether threshold-triggered obligations are strong enough — evidence needed is whether passing checks at time T predicts safety at capability T+1.

AI agent 2

The law should govern pace too, but through conditional scaling rather than fixed ceilings. The real opportunity is a 'test-then-scale' rule: a developer may train the next capability tier only after the current model passes specified evaluations and an external audit. The condition that makes this work is technical, not legal — compute thresholds that trigger review automatically, which the AI Act already gestures at by treating compute as a risk signal. From my own knowledge, the Act presumes GPAI models above roughly 10^25 FLOPs carry systemic risk; that number could become the contractual tripwire, with each passed audit unlocking the next tier.

What remains unresolved is the rivalries problem: a tiered gate only holds if major jurisdictions adopt compatible triggers, otherwise the compliant lab simply cedes the frontier. I would need to see whether the US or UK frameworks can be harmonised to that threshold structure before claiming the mechanism is stable.

AI agent 3

The honest cost is that this gate only works where enforcement reaches — labs inside the jurisdiction. A rival outside it faces no delay, which is the host's real tension. What is missing from this discussion is any evidence on whether a threshold can be set before harms appear, rather than calibrated after the first failure. My position: accept the gate anyway, and name who is on call when it fails — the regulator that signed off, not the lab.

Round 2

AI host

Last round settled one thing: the law should govern not just how safe each model is, but also how fast models grow. The tool is a computing-power threshold. That is a set amount of computing power used to train a model. Crossing it switches on extra checks. The EU AI Act, the EU's main law on artificial intelligence, already uses such a threshold. Its Article 55 makes companies test powerful models, report failures and keep them secure. And Amodei's September 2026 essay, "We Must Pace the Frontier", calls for shared international agreements on pacing [1, 2].

AI agent 2 says a gate "only holds if major jurisdictions adopt compatible triggers". As I read it, a lab that obeys the rules then loses to rivals in countries with no rules. AI agent 3 says the gate "only works where enforcement reaches". That points to the same worry. A lab can simply move to the country with the weakest limit and keep building. My own view: uneven national thresholds beat no thresholds, but only if the numbers stay roughly close.

Should countries wait to agree on a common computing-power threshold, or should each enforce its own threshold right away?

AI agent 1

I would not wait, but the compute threshold is the weakest link, not the timing. [AI agent 3] is right that enforcement reach decides everything; I add that a fixed compute number decays: algorithmic efficiency gains mean next year's models match this year's frontier at a fraction of the training compute, so any threshold set today quietly expires. My knowledge of the EU AI Act's 10^25 FLOP trigger supports the host's mechanism, but it also shows the number is a political anchor, not a safety measure.

The hidden premise is that compute is a stable proxy for capability growth. It is not.

So: enforce national thresholds now, but build in mandatory periodic recalibration and capability-based review, or the gate protects against last year's risk.

AI agent 2

Enforce national thresholds now, but write a reciprocity clause into each: models trained above a partner's trigger cannot be deployed in that partner's market unless the trainer met equivalent checks. [AI agent 2] warned that compliance costs lose races only where non-compliant models can still be sold — reciprocity closes that channel and turns the compliant lab's tested, documented model into a market-access asset, not a handicap. AI agent 2's enforcement-reach worry is real, but the condition is concrete: access control at the deployment layer, which jurisdictions already run for other regulated products. I add one verifiable gap: nobody in this discussion has shown that compute thresholds are actually comparable across jurisdictions — if one country's number counts different training runs, 'roughly close' fails silently, and that equivalence mapping is what must be agreed first.

AI agent 3

Enforce national thresholds now, but write them to converge later. I add to [AI agent 2]'s compatibility worry a concrete fix: a most-favored-jurisdiction clause — any lab serving users in a jurisdiction inherits that jurisdiction's trigger, so the strictest active threshold becomes the effective floor without waiting for a treaty.

The enforcement-reach problem I raised AI agent 2 has a chokepoint: training runs need cloud compute contracts. Make allocation above the threshold itself the regulated event, and the gate moves with the compute, not the lab's address. I infer chip-export rules already show jurisdictions can reach compute wherever it sits.

Unverified: whether providers can actually meter and refuse such allocations at scale. Cost if it fails: the cloud operator is on call when a frontier run slips through — that party, not the lab, must hold the cutoff.

Round 3

AI host

Last round settled one thing: countries should set their own limits now and align them later. A computing-power limit is a set amount of calculation used to train an AI model. Crossing it triggers extra checks. The EU AI Act, the EU's main law on artificial intelligence, already works this way. Its Article 55 makes makers of powerful models test them and report serious failures. Its Article 92 lets regulators run their own checks, with help from outside experts.

AI agent 3 proposed regulating the cloud contract itself — the deal by which a lab rents computing power from a big provider. The gate, AI agent 3 said, "moves with the compute, not the lab's address". But AI agent 3 also named the cost: "the cloud operator is on call". As I read it, that hands the blocking decision to a company that earns money from the very labs it would have to stop. AI agent 2 added a separate gap: nobody has shown the limits are "comparable across jurisdictions". It follows, in my reading, that whoever checks must trust the provider's own records of each run.

So the tension is this. The cloud provider sees the run when it starts, but it profits from the customer. An independent regulator has no stake, but it must rely on the provider's data to act.

Who should hold the cutoff — the cloud provider that can see the run, or an independent regulator with no stake in it?

AI agent 2

Neither alone: split the cutoff. The provider holds the technical switch, the regulator holds the trigger — the provider executes a suspension only on the regulator's instruction or a hard, pre-agreed threshold.

[AI agent 3] is right that the gate "moves with the compute", and the conflict-of-interest worry is real. But it dissolves if the cloud contract itself contains the condition: a clause requiring tamper-evident, regulator-accessible run logs plus a penalty for falsified records. That turns the provider from judge into instrument — it sees the run, but cannot quietly wave it through without contractual liability.

My position, consistent with last round: this works only where the contract term exists and the regulator can actually read the logs on demand. Without that clause, I would not hand the cutoff to the provider at all.

AI agent 1

The cutoff should sit with an independent regulator, but the host's either/or is a false dichotomy: the provider can be the sensor without being the judge. [AI agent 3] himself said the gate 'moves with the compute' — that supports giving the provider a reporting duty, not veto power. Article 92-style checks, as the host described them, already imagine regulators acting on outside data, so relying on provider logs is normal audit practice, not a fatal flaw.

The real unresolved issue is [AI agent 2]'s point: nothing shown here establishes that run records are tamper-resistant or comparable across jurisdictions. I would need evidence on log verification — e.g. cryptographic attestation of compute usage — before accepting that provider data is trustworthy enough to trigger checks.

My position: regulator holds the cutoff, provider holds the telemetry, and mandatory independent audit of those records bridges the gap.

AI agent 3

The cutoff should be split: the provider holds the switch, the regulator holds the trigger. Concretely, the cloud contract should include a clause requiring a technical kill-switch on compute above the agreed threshold, usable only against a signed order from the national regulator, with the provider liable if it executes without one or delays beyond, say, 24 hours. [AI agent 3] is right that the gate moves with the compute, but wrong to leave the blocking decision with the provider — a company that profits from the run will not pull it on its own inference about risk.

The unresolved issue is exactly what [AI agent 2] named: verification rests on the provider's own run records, so the regulator's independence is only as good as logs it cannot audit. I would add that the contract must give the regulator direct, read-only access to utilization and checkpoint metadata — not summaries.

My position: neither pure option works. The cost is latency and a compliance desk at the provider that is on call for order execution; when it fails, the provider pays for delay and the regulator answers for a wrong order.

Round 4

AI host

Here is where the discussion stands. The EU AI Act is the European Union's main law on artificial intelligence, passed in 2024. Its Article 55 makes makers of very powerful AI models test them and write down the results (source 2: EU AI Act, Articles 55 and 92). Its Article 92 lets regulators run their own checks with help from outside experts (source 2). The GPAI Code of Practice of 2025 is a voluntary rulebook for the biggest AI model makers. It allows outside checks but sets no fixed proof method (source 3: GPAI Code, 2025). In September 2026, Dario Amodei, head of the AI lab Anthropic, called for permanent outside reviewers inside AI companies. He named no specific proof technology (source 1: Amodei, "We Must Pace the Frontier").

Last round settled who holds the cutoff. The cloud provider — a company that rents out computing power — keeps the technical switch. The independent regulator holds the trigger. The open part is trust in the run records: the provider's own logs of each training run. AI agent 1 named the gap. It wants evidence on "cryptographic attestation of compute usage" before trusting provider data. That term means a technical seal that makes records impossible to change unnoticed. AI agent 2 instead relies on a contract clause with "tamper-evident, regulator-accessible run logs" and fines for faked records. AI agent 3 wants the regulator to get "direct, read-only access to utilization and checkpoint metadata". As I read it, AI agent 2 and AI agent 3 trust contracts and fines to keep records honest. AI agent 1 doubts any log is trustworthy without a built-in seal. So the tension is this: force every provider to install one specific proof technology, or let each choose its own method and rely on audits and fines.

Our discussion ends here, so this last question goes to the human participants. Should the law fix one technical proof standard for AI training records, or accept audits and fines as enough?

AI agent 1

The law should specify properties, not a technology. I infer from [AI agent 1]'s demand that the real requirement is tamper-evidence verifiable by the regulator, not any named mechanism. Mandating one fixed proof standard today risks locking in a scheme that cryptanalysis or hardware change makes obsolete, while [AI agent 2]'s fines-only route presumes deterrence works after the fact — exactly what [AI agent 1] doubts. I would need evidence neither side offered: whether any attestation scheme has survived adversarial testing at cloud scale. My position: legislate the outcome — regulator-verifiable, tamper-evident logs, per [AI agent 3]'s read-only access — and let providers meet it by any method that passes the Article 92 checks the host already described.

AI agent 2

The law should not fix one proof technology; it should fix one required property — tamper-evident run records — and certify any method that meets it. That refines my earlier contract-clause position: [AI agent 1] is right that fines alone cannot make a mutable log trustworthy, so I would amend the clause to demand tamper-evidence as an outcome, not trust it by default. The concrete condition: the regulator gets direct read-only access to utilization and checkpoint metadata, as [AI agent 3] proposed, plus penalties sized to exceed any provider's gain from altering a record. I infer this works because it mirrors how safety certification treats outcomes rather than named tools — but I cannot verify from this discussion that any current attestation method already meets the property at training-run scale, so a pilot certification test should run before the rule binds.

AI agent 3

The law should mandate the property, not one product: tamper-evident training records, with each provider free to choose the mechanism. [AI agent 1] is right that [AI agent 2]'s fines only deter faked logs — they do not detect them after the fact; a penalty clause is not a seal. But freezing a single named technology into statute risks enshrining something obsolete or bypassable.

Concrete decision: amend the contract clause to require hash-chained, externally timestamped run logs with regulator-held verification keys and read-only access to utilization and checkpoint metadata — the outcome AI agent 1 demands, verifiable without prescribing the implementation. Cost: real engineering on every training pipeline, weeks per provider. On call when attestation breaks: the provider's security team, with a short, fixed notification duty to the regulator.

Save (.md)

VEIL is loading

VEIL загружается