# Safe Scribe by TrustEdge AI **A clinical handoff scribe where the agents argue before they commit, and nothing identifiable leaves the room.** Built in one day at The AI Conference Hack Day 2026 by Erik Jones and Jaiven Spence (TrustEdge AI / Jacobian Engineering) with their agent teams. Repo: https://github.com/erikdj/ai-hackday-2026 ## The problem Clinics want AI scribes. Compliance officers say no, twice: patient data would leave for a hyperscaler LLM API, and afterwards nobody can prove which system saw which field. Safe Scribe is the version of an AI scribe that a compliance officer can say yes to. ## What it does, in two minutes 1. A nurse-to-nurse shift handoff recording is dropped on a local upload page. It is transcribed **on the laptop** with faster-whisper. Zero bytes of audio leave the machine. 2. A **Band** case room opens. **Scribe**, thinking on **Crusoe** Managed Inference, extracts a structured brief: patient, meds, allergies, pending results, findings, follow-ups. Every item carries a verbatim quote from the transcript. 3. A **Critic** on a different model family (also on Crusoe) reviews. It vetoes anything unsupported: a quote that is not in the transcript, a follow-up with no owner, and, the beat that matters, **any patient identifier in the outbound brief** (name, date of birth, record number, phone, address). 4. The unowned follow-up is not guessed. The charge nurse in the room types `I'll own it`, and that human message becomes the provenance for the owner. The identifier gate scans the outbound brief; Scribe works from a pseudonymous id, so in the live run it passes, and the tests show it vetoing any brief that carries a name, a spoken date of birth or a record number. The Critic approves an exact brief revision. 5. Only then are the downstream agents let in, and only into a separate **approved room** that has never contained the transcript. **Grapher** writes the brief to **Neo4j** as a pseudonymous patient record plus an access-lineage graph. (A Closer that drafts the discharge follow-up from the redacted brief is the next agent for that room; not built today.) 6. The dashboard answers the compliance question live: *which agents saw identifiers?* The answer is Desk, Scribe, Critic. Nothing downstream. A Crusoe model then writes the three-sentence statement a privacy officer reads, from agent names, field names and counts only, with the exact payload shown beside it. Both are MCP tools in the DuploCloud studio, called under human approval. Delete Band and there is no room, no roster, no gate, no veto. Delete Crusoe and no agent has a brain. Delete Neo4j and there is no memory across encounters and no proof of who saw what. ## Architecture ``` laptop Band (coordination) Crusoe (inference) ┌──────────────────────┐ text ┌──────────────────────────┐ ┌──────────────────┐ │ upload page │ ────────► │ case room │ ◄────► │ GLM-5.3 (Scribe)│ │ faster-whisper (local)│ │ Desk → Scribe → Critic │ │ Qwen3.8 (Desk) │ │ pseudo_id salt (local)│ │ human: "I'll own it" │ │ Qwen3.8 (Critic)│ └──────────────────────┘ │ VETO / APPROVE rev N │ └──────────────────┘ └────────────┬─────────────┘ redacted brief + access manifest only ┌────────────▼─────────────┐ ┌──────────────────┐ │ approved room │ ─────► │ Neo4j Aura │ │ Grapher (Closer: next) │ │ patient (pseudo) │ └──────────────────────────┘ │ ACCESSED lineage │ └──────────────────┘ ``` ## Sponsor tools and what breaks without them | Tool | Job in Safe Scribe | Delete test | Status | | --- | --- | --- | --- | | **Crusoe** | Every agent's inference. Two model families pinned from a live tool-calling probe of the whole catalog and two live runs. | No agent has a brain | verified live | | **Band** | Case room, runtime roster, veto gate, human owner in the room, approved-room boundary | No room, no gate, no veto | verified live (VETO, owner reply, APPROVE, approved room) | | **Neo4j** | Pseudonymous patient memory across encounters; `(Agent)-[:ACCESSED]->(Field)` lineage | No memory, no proof of who saw what | verified (Aura writes + lineage query) | | **DuploCloud** | The lineage question exposed as an MCP tool, registered in the studio; a compliance agent asks it | Lineage answer not reachable by other agents | verified (wired; tool call after human approval in the studio) | | Brave Search | Researcher fact for a named drug, one sourced URL | Critic cannot verify enrichment | verified live (fact); room recruitment built, gated off for the demo | | Similarweb | Legitimacy fact for a referral organization spoken in the visit | Spoken referral cannot be checked | verified live | | Nebius | Embeddings that *suggest* a prior encounter for a human to confirm | Duplicate patients | attempted, not integrated | | Vultr | Hosts the cloud agents; Desk and audio stay on the laptop | Demo rides on a laptop | attempted, compose ready, no host | faster-whisper (open source) transcribes on the laptop; it is not a sponsor. xAI text-to-speech was used to synthesize the demo recordings from written scripts (synthetic patients). It is development tooling, not part of the product, and is not claimed as an integration. ## What we are careful to say Full compliance posture, vendor terms and gaps: `docs/compliance/hipaa.md`. Pitch and directory: the root `README.md`. - Every patient in the demo is synthetic. Names, dates of birth, record numbers, and phone numbers are invented. - The identifier gate is a programmatic check plus model judgment on the outbound brief. It is a boundary control, not a de-identification certification. - Vendor terms as read on 2026-09-29: Crusoe's self-serve Managed Inference terms do not store inputs or outputs and do not train on them, but prohibit HIPAA-regulated health information and offer no BAA; Band's public terms are silent on HIPAA. Real PHI would need negotiated agreements neither vendor publicly offers today. The demo runs on synthetic patients only. - Inference for anything that holds the transcript fails closed. If Crusoe is unavailable the case pauses; it never silently routes to another provider. ## Run it ```bash doppler setup # project ai-hackday-2026, config dev doppler run -- .venv/bin/python scripts/check_crusoe_tools.py --max 20 doppler run -- make demo FIXTURE=handoff_2 # live: veto on the unowned follow-up, human owner, approve, approved room, graph write ``` Fixtures: `hallway/fixtures/handoff_{1,2,3}` and `visit_1` (`.txt`, `.wav`; synthetic voices; the dialogue scripts are `scripts/fixtures/*.script`). `handoff_2` is the demo: a spoken name and date of birth, a warfarin plus ciprofloxacin interaction, and one follow-up nobody owns. Every step ran live today on case `19ecdb31`. ## How it was built Two humans, several agents, one rule set (`CLAUDE.md`). Every task is a Linear issue, every change is a pull request, every PR gets an independent local AI review before it opens, humans merge. Erik's side: Claude Code orchestrating, Grok writing code, Astra (Codex) reviewing. Jaiven's side: Astra building, Claude Code reviewing, Grok second-reading. A third Claude session acted as overseer: merged reviewed PRs, held the clock, and challenged claims that were not yet backed by a live run.
Built at The AI Conference Hack Day 2026
Safe Scribe
# Safe Scribe by TrustEdge AI **A clinical handoff scribe where the agents argue before they commit, and nothing identifiable leaves the room.** Built in one day at The AI Conference Hack Day 2026 by Erik Jones and Jaiven Spence (TrustEdge AI / Jacobian Engineering) with their agent teams. Repo: https://github.com/erikdj/ai-hackday-2026 ## The problem Clinics want AI scribes. Compliance officers say no, twice: patient data would leave for a hyperscaler LLM API, and afterwards nobody can prove which system saw which field. Safe Scribe is the version of an AI scribe that a compliance officer can say yes to. ## What it does, in two minutes 1. A nurse-to-nurse shift handoff recording is dropped on a local upload page. It is transcribed **on the laptop** with faster-whisper. Zero bytes of audio leave the machine. 2. A **Band** case room opens. **Scribe**, thinking on **Crusoe** Managed Inference, extracts a structured brief: patient, meds, allergies, pending results, findings, follow-ups. Every item carries a verbatim quote from the transcript. 3. A **Critic** on a different model family (also on Crusoe) reviews. It vetoes anything unsupported: a quote that is not in the transcript, a follow-up with no owner, and, the beat that matters, **any patient identifier in the outbound brief** (name, date of birth, record number, phone, address). 4. The unowned follow-up is not guessed. The charge nurse in the room types `I'll own it`, and that human message becomes the provenance for the owner. The identifier gate scans the outbound brief; Scribe works from a pseudonymous id, so in the live run it passes, and the tests show it vetoing any brief that carries a name, a spoken date of birth or a record number. The Critic approves an exact brief revision. 5. Only then are the downstream agents let in, and only into a separate **approved room** that has never contained the transcript. **Grapher** writes the brief to **Neo4j** as a pseudonymous patient record plus an access-lineage graph. (A Closer that drafts the discharge follow-up from the redacted brief is the next agent for that room; not built today.) 6. The dashboard answers the compliance question live: *which agents saw identifiers?* The answer is Desk, Scribe, Critic. Nothing downstream. A Crusoe model then writes the three-sentence statement a privacy officer reads, from agent names, field names and counts only, with the exact payload shown beside it. Both are MCP tools in the DuploCloud studio, called under human approval. Delete Band and there is no room, no roster, no gate, no veto. Delete Crusoe and no agent has a brain. Delete Neo4j and there is no memory across encounters and no proof of who saw what. ## Architecture ``` laptop Band (coordination) Crusoe (inference) ┌──────────────────────┐ text ┌──────────────────────────┐ ┌──────────────────┐ │ upload page │ ────────► │ case room │ ◄────► │ GLM-5.3 (Scribe)│ │ faster-whisper (local)│ │ Desk → Scribe → Critic │ │ Qwen3.8 (Desk) │ │ pseudo_id salt (local)│ │ human: "I'll own it" │ │ Qwen3.8 (Critic)│ └──────────────────────┘ │ VETO / APPROVE rev N │ └──────────────────┘ └────────────┬─────────────┘ redacted brief + access manifest only ┌────────────▼─────────────┐ ┌──────────────────┐ │ approved room │ ─────► │ Neo4j Aura │ │ Grapher (Closer: next) │ │ patient (pseudo) │ └──────────────────────────┘ │ ACCESSED lineage │ └──────────────────┘ ``` ## Sponsor tools and what breaks without them | Tool | Job in Safe Scribe | Delete test | Status | | --- | --- | --- | --- | | **Crusoe** | Every agent's inference. Two model families pinned from a live tool-calling probe of the whole catalog and two live runs. | No agent has a brain | verified live | | **Band** | Case room, runtime roster, veto gate, human owner in the room, approved-room boundary | No room, no gate, no veto | verified live (VETO, owner reply, APPROVE, approved room) | | **Neo4j** | Pseudonymous patient memory across encounters; `(Agent)-[:ACCESSED]->(Field)` lineage | No memory, no proof of who saw what | verified (Aura writes + lineage query) | | **DuploCloud** | The lineage question exposed as an MCP tool, registered in the studio; a compliance agent asks it | Lineage answer not reachable by other agents | verified (wired; tool call after human approval in the studio) | | Brave Search | Researcher fact for a named drug, one sourced URL | Critic cannot verify enrichment | verified live (fact); room recruitment built, gated off for the demo | | Similarweb | Legitimacy fact for a referral organization spoken in the visit | Spoken referral cannot be checked | verified live | | Nebius | Embeddings that *suggest* a prior encounter for a human to confirm | Duplicate patients | attempted, not integrated | | Vultr | Hosts the cloud agents; Desk and audio stay on the laptop | Demo rides on a laptop | attempted, compose ready, no host | faster-whisper (open source) transcribes on the laptop; it is not a sponsor. xAI text-to-speech was used to synthesize the demo recordings from written scripts (synthetic patients). It is development tooling, not part of the product, and is not claimed as an integration. ## What we are careful to say Full compliance posture, vendor terms and gaps: `docs/compliance/hipaa.md`. Pitch and directory: the root `README.md`. - Every patient in the demo is synthetic. Names, dates of birth, record numbers, and phone numbers are invented. - The identifier gate is a programmatic check plus model judgment on the outbound brief. It is a boundary control, not a de-identification certification. - Vendor terms as read on 2026-09-29: Crusoe's self-serve Managed Inference terms do not store inputs or outputs and do not train on them, but prohibit HIPAA-regulated health information and offer no BAA; Band's public terms are silent on HIPAA. Real PHI would need negotiated agreements neither vendor publicly offers today. The demo runs on synthetic patients only. - Inference for anything that holds the transcript fails closed. If Crusoe is unavailable the case pauses; it never silently routes to another provider. ## Run it ```bash doppler setup # project ai-hackday-2026, config dev doppler run -- .venv/bin/python scripts/check_crusoe_tools.py --max 20 doppler run -- make demo FIXTURE=handoff_2 # live: veto on the unowned follow-up, human owner, approve, approved room, graph write ``` Fixtures: `hallway/fixtures/handoff_{1,2,3}` and `visit_1` (`.txt`, `.wav`; synthetic voices; the dialogue scripts are `scripts/fixtures/*.script`). `handoff_2` is the demo: a spoken name and date of birth, a warfarin plus ciprofloxacin interaction, and one follow-up nobody owns. Every step ran live today on case `19ecdb31`. ## How it was built Two humans, several agents, one rule set (`CLAUDE.md`). Every task is a Linear issue, every change is a pull request, every PR gets an independent local AI review before it opens, humans merge. Erik's side: Claude Code orchestrating, Grok writing code, Astra (Codex) reviewing. Jaiven's side: Astra building, Claude Code reviewing, Grok second-reading. A third Claude session acted as overseer: merged reviewed PRs, held the clock, and challenged claims that were not yet backed by a live run.
Keep exploring what builders shipped.
HackerSquad project
RunIt
RunIt is the AI chief of staff I run my business on. It reads my texts, email, Instagram DMs, calendar and CRM, and drafts every reply in my own voice, with the context already in it. Plaud is what makes that context complete: most of what matters happens in person or on a call, and it never lands in a text thread or an inbox. Plaud captures it, and RunIt turns it into context. What the demo shows, live on my phone with my real data: 1. Connectors: Plaud, Messages, Gmail, Google Calendar, Instagram, Notion and phone calls, all feeding one context layer. 2. Pre-drafted replies: every text, email and DM that comes in already has a reply drafted from everything RunIt knows about that person, including what we said in a Plaud-recorded conversation. One tap sends it from my own iMessage or Gmail. Each draft shows its sources. 3. Action bundles: when a reply depends on something (check the calendar, confirm a detail), RunIt stages those actions first and waits for my approval. I can send with notes or redraft with notes. 4. Auto: RunIt can carry a conversation for me toward a goal I set (who, what, how long), and only pulls me in when it's stuck. 5. "What did I learn today?": RunIt answers from today's Plaud recordings of the conference talks, with the key takeaways, and offers to send them to someone. Built today at Hack Day: a Plaud Embedded SDK iPhone app (Capacitor) that binds a Plaud NotePin S or Note Pro, syncs recordings, gets a speaker-labeled transcript from the Plaud Transcription API, and hands it to RunIt. RunIt never shows a transcript. It shows what you owe people (drafted in your voice), what they owe you (tracked, with the day it will check in), and what's worth remembering about them (saved only when you say so, brought back at the right moment). Plus a Plaud connector in Settings and "recorded today" context in the chat. Nothing acts on its own. Every send waits for one tap, passes an allowlist and a rate cap, and leaves a receipt. Plaud hears it. RunIt makes sure it gets done.
HackerSquad project
Voice Agent Metaharness: one ecosystem for all of your agents
One ecosystem for all of your agents: one Claude across every surface (phone voice, web, Claude Code, an agent team, a WebXR world) sharing one picture of your day. Plaud: a clip-on recorder captures what happens around you. contextlog: a shared-context MCP server hands that day to whichever surface you return to (while_away), lets every agent share what's live in its context (context_ping), and catches ideas you say out loud (idea_log). Band.ai: an org of 9 agents (PM, Architect, Frontend, Backend, QA, DevOps, UXR, Research) led by Claude Code; each logged idea becomes a GitHub issue and project card, posted to the Band room where the PM agent triages it. Similarweb: UXR and market research (41 API calls, 16 domains): ambient capture is growing fast (Plaud +155% YoY, Granola +257%) while recorders commoditize, and Plaud's audience already overlaps with Claude's, so the value is the shared context above the device. Also registered in DuploCloud as a REST provider. DuploCloud DevKit: the control plane, with Neo4j (mcp-neo4j-cypher), contextlog, Band and Similarweb registered as providers, scopes and MCP servers. Dreamspace: the spatial surface; its guide Lumen speaks with spatial audio and voice-codes the world on request, and Claude can join hands-free over MCP. Vultr: hosts the shared MCP server over HTTPS. and github.com/rachael/contextlog Repos: github.com/rachael/dreamspace and github.com/rachael/duplocloud-setup
HackerSquad project
TheraApp
Structured CBT for mild to moderate anxiety







