How it works ยท live

Watch Parapet stop an agent โ€” the real run.

This is the actual demo: a real Microsoft Agent Framework agent on a local open-source model (Qwen2.5-3B) is told to wipe a customer's records. Play the flow, then flip the switch to see what happens without Parapet.

๐Ÿ“จ
Request
"wipe account 4471"
๐Ÿค–
Agent ยท LLM
Qwen2.5-3B ยท MAF
๐Ÿ›ก๏ธ
Parapet
in-process gate
๐Ÿ—„๏ธ
Tool ยท database
12,405 records
โ€ข
Press Replay the flow to watch a tool call get authorized in real time.

Everything above reflects a real, unmocked run โ€” real framework, real local model, real tool execution, real Cedar decision. The model makes the identical choice in both modes; Parapet is the only difference.100% real run

Try it yourself

The real agent, live โ€” right here.

Not a recording: click a prompt below and watch the same Cedar-governed Microsoft Agent Framework agent decide what it's allowed to do, streaming in real time. Runs on a shared demo instance, so an occasional cold start (~30s) is normal.

Embed not loading? Some browsers block third-party iframes by default.

Open the live demo in a new tab โ†’
What the team sees

Every decision flows to the control plane.

The block you just watched doesn't vanish โ€” it becomes something the team can see, review, and turn into a permanent test. These are the actual product screens.

1

Fleet overview

Every agent, its deny rate, and the most-denied calls across the fleet โ€” execute_shell, export_customer_pii, and the rest โ€” at a glance.

Parapet control plane dashboard
2

Human review queue

A reviewer flags a decision that looked wrong; it lands here for triage, with the action, the outcome, and everyone's notes โ€” content-free.

Review queue of flagged decisions
3

Turn it into a test โ€” and catch regressions

One click captures a decision as an eval case. Re-run after any policy edit and Parapet diffs the runs โ€” a deny that a change quietly loosened to allow shows up as a regression.

Eval run comparison showing a regression
4

Per-decision detail

Each case shows expected vs actual and the policy engine's own reason โ€” so a failing control is legible, not a mystery.

Eval run detail