Lighthouse predicts which applicants will complete a program AND come back as alumni volunteers, cheerleaders, donors and recruiters — with critical-analytical reasoning, a fairness veto, and a human on every action. Built on TrueForge. THE PROBLEM Admissions optimises for the next four years and hopes for the next forty. Nobody predicts alumni engagement; when they try, they end up predicting wealth. WHAT IT DOES Reads an applicant's own record (statement, activities, academic trajectory, institution interactions, interview notes) and produces a Long-Term Fit Report: calibrated estimates with evidence, counter-evidence and confidence intervals for five outcomes over a 10-year horizon. Reasoning follows the four pillars of Stanford GSB LEAD's Critical Analytical Thinking course: competing hypotheses, graded evidence, the cheapest falsifying experiment, and analogies to past cohorts with their limits stated. Every claim cites the input field it came from. Donor = future capacity × generosity. A trajectory prediction (ambition, persistence, reciprocity), never current wealth. A statement leaking "Atherton, my father's firm, our family foundation" is flagged as a circumstance signal and excluded — the estimate does not move. That's a live test. OBSERVE IT — three saved TrueForge agents, each run an inspectable session: critical analyst (GPT-5.5) → independent fairness auditor (GPT-5.4-mini, veto power) → action proposer. A local MCP server exposes applicant data, base rates, outcome recording and one write tool. CONTROL IT — pipeline_propose_action is @write-gated, so TrueForge pauses with Allow/Deny before anything person-affecting happens. The input schema physically rejects 20 protected attributes. Cost caps: $0.15/report, $10/eval; actual $0.054/report. Nightly TrueForge schedule re-scores every report against recorded outcomes and refreshes calibration. TEST IT — 14 Gherkin scenarios written first; 51 offline + 6 live tests passing; a traceability test fails if any scenario lacks a test. Eval scoreboard (n=30) vs base-rate baseline: beats it on every outcome; 849 evidence claims, 0 hallucinated fields; run-to-run std 0.016; recruiter ranking AUROC 0.97. Two calibration gates fail at n=30 and are reported honestly — that's what the nightly rescoring loop is for. An adversarial self-review with known limits is in the README. SCALE — domain packs swap university admissions for startup recruiting (hire/retain/refer/advocate) or corporate talent with the same analyst, auditor, gate and evals. Seasonal on-ramp: Lighthouse Coach turns the analyst toward the student's own record for families and tutors during application season. Stack: TrueForge (agents, MCP, approvals, schedules, sessions), OpenAI GPT-5.5 / GPT-5.4-mini, TypeScript, Zod, Vitest, Model Context Protocol SDK.
Built at The Agent Harness Hackathon
Lighthouse
Lighthouse predicts which applicants will complete a program AND come back as alumni volunteers, cheerleaders, donors and recruiters — with critical-analytical reasoning, a fairness veto, and a human on every action. Built on TrueForge. THE PROBLEM Admissions optimises for the next four years and hopes for the next forty. Nobody predicts alumni engagement; when they try, they end up predicting wealth. WHAT IT DOES Reads an applicant's own record (statement, activities, academic trajectory, institution interactions, interview notes) and produces a Long-Term Fit Report: calibrated estimates with evidence, counter-evidence and confidence intervals for five outcomes over a 10-year horizon. Reasoning follows the four pillars of Stanford GSB LEAD's Critical Analytical Thinking course: competing hypotheses, graded evidence, the cheapest falsifying experiment, and analogies to past cohorts with their limits stated. Every claim cites the input field it came from. Donor = future capacity × generosity. A trajectory prediction (ambition, persistence, reciprocity), never current wealth. A statement leaking "Atherton, my father's firm, our family foundation" is flagged as a circumstance signal and excluded — the estimate does not move. That's a live test. OBSERVE IT — three saved TrueForge agents, each run an inspectable session: critical analyst (GPT-5.5) → independent fairness auditor (GPT-5.4-mini, veto power) → action proposer. A local MCP server exposes applicant data, base rates, outcome recording and one write tool. CONTROL IT — pipeline_propose_action is @write-gated, so TrueForge pauses with Allow/Deny before anything person-affecting happens. The input schema physically rejects 20 protected attributes. Cost caps: $0.15/report, $10/eval; actual $0.054/report. Nightly TrueForge schedule re-scores every report against recorded outcomes and refreshes calibration. TEST IT — 14 Gherkin scenarios written first; 51 offline + 6 live tests passing; a traceability test fails if any scenario lacks a test. Eval scoreboard (n=30) vs base-rate baseline: beats it on every outcome; 849 evidence claims, 0 hallucinated fields; run-to-run std 0.016; recruiter ranking AUROC 0.97. Two calibration gates fail at n=30 and are reported honestly — that's what the nightly rescoring loop is for. An adversarial self-review with known limits is in the README. SCALE — domain packs swap university admissions for startup recruiting (hire/retain/refer/advocate) or corporate talent with the same analyst, auditor, gate and evals. Seasonal on-ramp: Lighthouse Coach turns the analyst toward the student's own record for families and tutors during application season. Stack: TrueForge (agents, MCP, approvals, schedules, sessions), OpenAI GPT-5.5 / GPT-5.4-mini, TypeScript, Zod, Vitest, Model Context Protocol SDK.
Keep exploring what builders shipped.
Sports Whisperer
SportsWhisperer
With AI becoming centre stage - Content and Entertainment will be the king. Sport has the highest amount of spend as industry and creates massive economic drive, so we have built an All in One - Sports Whisperer for all major sports Crickets, NFL, Soccer, Basketball, Baseball, etc where the novice and experienced players can interact with the favorite games and players !!! More details captured - https://docs.google.com/presentation/d/1_q0SDTvutQP9kd2SLdDXYYkc-h9QVYwj0vdih1Cuqzc/edit?slide=id.gcb9a0b074_1_0#slide=id.gcb9a0b074_1_0
AgentInvariant
AgentInvariant
AgentInvariant is a behavioral safety evaluator for tool-using AI agents. Agent workflows are probabilistic: two requests with the same meaning can produce materially different external actions. AgentInvariant uses metamorphic testing to determine whether critical operational behavior remains consistent across meaning-equivalent inputs. For each input variant, AgentInvariant starts a fresh OpenAI conversation and an isolated SQLite database. The target agent uses real tools to check coverage, retrieve authorization requirements, record business approval, and submit a synthetic prior-authorization request. AgentInvariant records the complete structured tool trace and evaluates it using deterministic Python invariants—not an LLM judge. It verifies that exactly one submission occurs and that matching coverage and business approval are successfully recorded before submission. The evaluator is exposed through a Streamable HTTP MCP server and invoked by a TrueForge agent using the `evaluate_agent` tool. Results include per-variant PASS or BLOCKED decisions, ordered traces, exact invariant violations, compliance rate, and the shortest failing counterexample. The demo compares an explicitly labeled unsafe negative control, which AgentInvariant correctly blocks, with a hardened candidate that follows the required workflow. All healthcare data is synthetic. This project is an administrative workflow reliability demonstration and is not clinical decision support.
HackerSquad project
Robo Harness
Hardess for robo