Skip to content
Stephen Binge
All work

Enterprise SaaS, built from scratch

Enterprise AI recruitment platform, built from scratch

An enterprise SaaS recruitment platform, built from scratch, that scores CVs and runs first-round voice interviews, with every decision left to a human and designed for the EU AI Act's high-risk rules.

Role
Technical lead and architect (shareholder)
Client
Recruitment SaaS company (anonymised)

The problem

Recruiters spend most of their time on the first pass: reading CVs against a brief, scheduling and running screening calls, and writing up notes. This platform does that first pass with AI so the recruiter's time goes on the candidates worth talking to.

Because this is HR, it falls under the EU AI Act's high-risk category. That shaped the design more than any model choice did.

What they already had

Nothing to build on. This was a new product, so the whole platform was built from scratch as an enterprise SaaS: the recruiter-facing app, the AI scoring and interview services, the human review workflow and the infrastructure behind them. Most of the work on this site starts from systems a client already runs. This one is here because sometimes the right answer is a clean build, and it shows what that looks like when it has to stand up to regulation.

The constraint that drove the design

The AI never makes a hiring decision. It produces research, scores against a published rubric and recommendations for a human reviewer, and every output is auditable. That means structured outputs rather than free text, a record of what the model saw and said, and a workflow where a person signs off.

What I built

  • CV assessment against a rubric of 4 categories and 13 scoring components, returned as structured JSON so scores can be compared, audited and re-run.
  • Real-time AI voice interviews on the web and over the phone: LiveKit Agents for the real-time session, Deepgram for speech-to-text, Claude for the conversation, text-to-speech back to the candidate, Twilio for the phone leg.
  • Video behavioural analysis as supporting evidence for reviewers.
  • Candidate flag and investigation workflow with an AI investigation assistant that gathers context for the reviewer. The reviewer decides.
  • An LLM evaluation harness to choose the scoring model: golden datasets and structured-output scoring across Claude, OpenAI and Gemini models, so model swaps are a measured decision.
  • Durable background workflows (Vercel Workflow) for long-running jobs, with retries and resumption.

What I deliberately didn't build

  • A model that outputs "hire" or "reject".
  • Free-text verdicts. Everything scored is structured and traceable to the rubric.
  • Custom real-time infrastructure. LiveKit, Deepgram and Twilio are production services; the work went into the agent logic and the review workflow.
For the technical readerUnder the hood: production details, architecture and stack

Architecture

  1. CandidateWeb or phone (Twilio)
  2. Voice agentLiveKit + Deepgram + Claude
  3. ScoringRubric, structured JSON, evals
  4. ReviewerHuman decision, audit trail
Every AI output lands in front of a reviewer. Nothing flows from the model straight to a decision.

Stack

Next.js, TypeScript, Supabase, LiveKit Agents, Python, Claude, OpenAI, Gemini, Deepgram, Twilio, AWS S3, Vercel, Vercel Workflow.

Fifteen minutes, free, and you’ll know whether it’s worth doing.

Bring the problem, or the idea. No sales pitch, and no obligation either way. You’ll leave with a straight answer on the simplest route and roughly what it would take.

Prefer email? me@stephenbinge.com