Page 2 — Where This Case Study Sits
What RansahAI is
RansahAI is an AI-native talent acquisition platform for B2B teams hiring at scale. Our flagship product, Elegant V1, runs the high-volume, low-judgment portion of the recruiting funnel — structured first-round screening interviews — and hands the structured output to a human recruiter for the high-judgment decisions that follow.
We sell a credit-based subscription across three tiers (Starter / Pro / Business) priced by interview volume. Token-level unit economics are tracked through Stripe Billing and SQL, so we can hold gross margin even as model costs shift.
What this case study is about
The case study covers one slice of Elegant V1: the candidate-facing AI interview experience. Specifically: how a job applicant — often nervous, often skeptical, often interviewing while their current employer doesn't know — experiences being interviewed by an AI for 15–25 minutes, and how we designed that experience to produce both (a) a fair, structured signal for the recruiter and (b) a candidate experience that does not poison the employer brand.
The recruiter-side dashboard, the orchestration backend (LangGraph + n8n), the eval infrastructure (DeepEval + Ragas + DSPy), and the pricing architecture are all out of scope for this document — each deserves its own case study.
Why this slice
Three reasons.
First, the candidate-side experience is where most AI recruiting tools fail publicly. HireVue, the category's most well-known incumbent, has been the subject of ongoing public criticism for the candidate experience of its asynchronous video interviews — including a 2019 EPIC complaint to the FTC and material reputational damage among candidate-facing reviews.
Second, it's where the product's regulatory exposure is highest. NYC's Local Law 144 (the Automated Employment Decision Tool law, in force since July 2023), Illinois's AI Video Interview Act, the EU AI Act's classification of recruitment AI as high-risk, and India's DPDP Act all impose obligations that touch the candidate's experience directly — disclosure, consent, the right to a human alternative, bias audits. Decisions made here in Phase I locked in our regulatory posture for years.
Third, it's where a young company can compete with incumbents. HireVue and Paradox have ATS-distribution moats that we cannot match in Year 1. They also have legacy candidate-experience problems they can't easily fix without breaking enterprise contracts. The candidate experience was a place where being small and new was a real advantage.
Page 3 — The Problem
The recruiter's reality
The recruiter buying RansahAI lives this loop: a single job posting receives 200–400 applications in its first week. The recruiter has roughly 6–8 seconds of attention per resume in the first pass. Most resumes are filtered by ATS keyword match, which means a strong candidate with non-standard phrasing is silently lost, and a weak candidate with the right keywords is silently advanced.
To compensate, recruiters either (a) outsource to agencies at 15–25% of first-year salary, (b) invest in a heavy enterprise screening tool like HireVue (~$35K–$300K/year, deep ATS integration required), or (c) skip first-round screening entirely and burn hiring-manager time on full-length interviews with weak candidates.
The candidate's reality
What gets less airtime in vendor decks is the candidate's version of the same problem. From the survey and interview research described later in this case study:
- The median candidate in our sample applied to 27 roles in the previous 90 days . Of those, they heard back from 4. Of those 4, 1 led to an interview. Application fatigue is the dominant emotional state.
- When candidates do receive an AI screening invitation, the dominant first reaction is not curiosity. It is suspicion — "is this real, or is the company too lazy to talk to me?"
- The candidates with the most to gain from a fair AI interview (those whose resumes don't keyword-match well — non-traditional career paths, international candidates, career switchers) are also the candidates most likely to distrust the AI screening.
The strategic question
“How might we design an AI-led first-round interview that produces a structured, defensible hiring signal for the recruiter, while leaving the candidate feeling more respected than they would have been by the alternative — silence, an ATS rejection, or a 20-minute Zoom with a junior recruiter who skimmed their resume on the way to the call?
Reading that HMW back: it deliberately doesn't optimize for "candidates love AI interviews." That bar is unreachable in 2024–2025 — the cultural prior is too negative. The bar we set was lower and harder: more respected than the realistic alternative. Most AI interview vendors set the bar at "candidates tolerate it," and tolerate is not enough.
We are positioned in the Discover/Define half of the Double Diamond for the duration of this case study. The Develop/Deliver work is documented separately.
Page 4 — Success Metrics
Picking a North Star that survives Goodhart's Law
The obvious vendor metric — number of interviews completed per month — is what we charge for. It's also exactly the metric that, if we optimized it directly, would degrade the product. A vendor incentivized to maximize interview volume will tolerate poor candidate experience, will tolerate low-quality questions, will tolerate weak signal-to-noise on the resulting reports, because none of those degrade the volume metric.
So our North Star is not interviews completed. It is:
“Recruiter-Action Rate (RAR): the percentage of completed interviews that result in the recruiter taking a defined next-stage action (advance, reject with documented reason, or shortlist) within 72 hours of completion.
This works because RAR fails when any of three things go wrong: candidates don't complete (no interview to act on), the recruiter doesn't trust the report (won't act on it), or the recruiter ignores the platform (action happens elsewhere and doesn't get logged). It is the smallest single number that catches the most failure modes.
Supporting metrics, with calibrated targets
Targets are based on our internal pilot data plus published benchmarks. They are targets, not aspirations — meaning we're prepared to ship at these levels, not "100% success" wishes.
| Layer | Metric | Target | Why this number |
|---|---|---|---|
| Funnel | Candidate completion rate (started → completed) | ≥ 78% | Asynchronous video tools cluster around 60–70% completion; we're targeting a meaningful improvement, not a magical one. |
| Funnel | Time-to-completion (median) | ≤ 22 minutes | Recruiter requirement: under the 30-minute "Zoom slot" mental model. |
| Trust | Candidate CSAT post-interview (1–5) | ≥ 3.8 | Industry CSAT for AI interviews appears to sit around 2.5–3.0. 4.5+ is unrealistic given the cultural prior. |
| Trust | Net Promoter Score among candidates | ≥ 0 (i.e., not net-negative) | An honest "we're not making things worse" target. |
| Signal quality | Inter-rater reliability between AI report and recruiter post-interview rating | Cohen's κ ≥ 0.55 (moderate agreement) | A stretch from typical resume-screening reliability of ~0.20–0.30. |
| Equity | Demographic completion gap (max group vs. min group) | ≤ 5 percentage points | NYC AEDT requires bias audits; we hold ourselves to a tighter internal bar than the audit threshold. |
| Commercial | Recruiter-Action Rate (North Star) | ≥ 65% | Anchored to the 72-hour SLA we promise enterprise pilots. |
What we explicitly chose not to measure as a goal
Vanity metrics we are not optimizing for, even though they look good in decks:
- "Candidates who said they preferred the AI interview to a human one." We don't believe this number can be honestly produced from a candidate population that has just spent 20 minutes talking to an AI; the answer is too contaminated by social desirability bias to be useful.
- "Diversity of candidates advanced." This metric, in isolation, is dangerous — it can be moved by reducing standards across the board, which helps no one. We track demographic completion and advancement gaps separately, in audit form, not as a target to hit.
Page 5 — Secondary Research
Market context
The AI talent acquisition market sits at the intersection of three forces:
(1) Recruiter overload is structural. Average tech requisitions receive 3–4× the application volume they did in 2019. The corresponding recruiter headcount has not scaled. The math forces some form of automation; the only question is which form.
(2) Candidate trust in AI hiring is at a historic low. Public coverage of bias incidents — the Amazon resume-screening tool that learned gender bias and was scrapped in 2018, ongoing concerns about facial-analysis features in video interviews, the 2023 EEOC settlement with iTutorGroup over age-discriminatory AI screening — has primed the cultural conversation negatively. Anything we ship is being read against that context.
(3) Regulatory tightening is real and accelerating.
| Jurisdiction | Regulation | What it requires |
|---|---|---|
| New York City | Local Law 144 (AEDT) | Annual bias audit, candidate disclosure, alternative process available |
| Illinois | AI Video Interview Act | Consent, disclosure, deletion rights |
| Maryland | Facial recognition in interviews | Banned without consent |
| EU | AI Act (in force from 2024–2026 phased) | Recruitment AI classified as high-risk; conformity assessments, human oversight, transparency |
| India | DPDP Act 2023 | Consent, purpose limitation, data principal rights |
| GDPR (EU) | Article 22 | Right not to be subject to fully automated decisions with significant effects |
This is not legal background music. It defined product decisions: every Elegant V1 candidate flow includes explicit AI disclosure, recorded consent, an opt-out path with no penalty, and a non-AI fallback. Our GCP-to-AWS migration was driven primarily by EU data residency obligations under GDPR.
Real-world benchmarks worth borrowing
Three external data points anchored our targets:
- Greenhouse and Lever benchmark data on time-to-hire and stage-conversion rates, used to justify the ≥65% Recruiter-Action Rate target.
- HireVue's published completion rates in their own marketing (claimed ~85% on assessments) — we discounted this to ~68% as the realistic baseline given selection effects in vendor-published numbers.
- Glassdoor candidate-experience reviews for AI interview tools — a sample of 200 recent reviews showed an average rating of 1.9/5 with the most common complaint being "I felt like I was talking to a wall." This is the qualitative bar we were trying to clear.
Page 6 — Competitive Teardown
A quick orientation. The market splits into four archetypes; we are deliberately positioned in the fourth.
Archetype 1 — Asynchronous video screening (HireVue, modern equivalents)
Candidate records video answers to pre-set questions; AI scores or surfaces clips for recruiters. Strengths: scale, ATS distribution. Weaknesses: candidate experience is widely disliked; candidate sees no other person and has no chance to clarify; facial-analysis features have created regulatory and reputational risk. Recent versions have removed facial analysis but the brand association persists. Pricing: enterprise, $35K–$300K/year.
Archetype 2 — Conversational chatbot (Paradox / Olivia, Mya, others)
Candidate interacts in chat, primarily for sourcing/scheduling rather than substantive interviewing. Strengths: low friction, fast adoption, strong scheduling features. Weaknesses: the conversational layer is shallow — most are scripted decision trees with LLM polish rather than genuine probing interviews. Doesn't replace a screening interview, supplements one. Pricing: mid-market enterprise, $25K–$100K/year.
Archetype 3 — Marketplace (Mercor and similar)
Vendor pre-screens candidates and presents a shortlist; charges commission on hire (often 25% of first-year compensation). Strengths: candidates are pre-vetted; employer doesn't run a funnel. Weaknesses: employer cedes funnel ownership and brand control; works best for narrow elite-talent verticals. Pricing: commission, not subscription.
Archetype 4 — AI + Human structured first-round (where Elegant V1 sits)
AI conducts a real conversational interview (voice-first, multi-turn, adaptive follow-ups), produces a structured report with full transcripts, hands off to a human recruiter who decides. Explicitly not a replacement for human judgment in the hire decision. Pricing: credit-based subscription, $0.X–$X per interview at the unit level [REAL — pricing tiers public].
The "AI + Human" framing is a product decision, not just marketing copy. Three implications follow from it:
- The AI doesn't decide. The output is structured evidence, not a hire/no-hire score. This sidesteps the hardest interpretation of EU AI Act / GDPR Article 22 (no fully automated decisions with significant effects) and makes the product easier to explain to a Chief People Officer who is afraid of headlines.
- The transcript is the artifact, not the score. The recruiter sees a full, attributed transcript with annotations — not a black-box "candidate fit score." This was a costly choice (transcripts are expensive to host and audit-review) and a deliberate one.
- Pricing is per-interview, not per-hire. Per-hire pricing aligns vendor incentives with hiring volume; per-interview pricing aligns vendor incentives with throughput, which is closer to the recruiter's actual job.
Where we are weak
I want to be specific about this rather than handwave. Compared to HireVue we have no ATS distribution, no enterprise security certifications yet, no track record. Compared to Paradox we have no scheduling layer in V1. Compared to Mercor we have no candidate supply — the customer has to bring their own funnel. These are real disadvantages and they constrain who we can sell to in Phase I (~50–500 employee tech and services companies, where buying decisions are faster and ATS lock-in is weaker).
Page 7 — Primary Research
Sample design
We ran two studies with deliberately separated populations to avoid the contamination that happens when interview participants also fill out the survey.
Study 1: Semi-structured candidate interviews. n = 24, 45–60 minutes each, conducted October–November 2024. Recruited via three channels to mitigate self-selection: (a) candidates who had completed a real Elegant V1 pilot interview (n=10), (b) candidates who had been invited but did not complete (n=8), (c) candidates who had never used an AI interview product but were active job seekers (n=6). The "did not complete" segment was the most valuable and the hardest to recruit; we paid a higher honorarium to get them.
Study 2: Quantitative survey. n = 184 active job seekers in tech and adjacent roles. Median tenure: 4 years. Geographic split: India 41%, US 22%, EU 18%, SEA 12%, other 7% . Survey was structured to force tradeoffs (rank these, choose one) rather than rate-on-a-scale, because rate-on-a-scale produces inflated agreement.
What I'm not claiming. With n=24 qualitative + n=184 survey, we have directional signal, not statistical certainty. The personas in the next section are hypotheses calibrated against this sample, not proven segments. Confidence intervals on any single survey number are approximately ±7 percentage points.
What recruiters told us (separate study, n=12 hiring managers and recruiters)
I'm including the recruiter findings briefly here because the candidate experience can't be designed without them:
- Recruiters universally distrusted single-score outputs. "If you give me a 7.4/10 I will spend the same time second-guessing it as I would have spent doing the screen myself" was the modal sentiment.
- Recruiters wanted exact quote evidence for any claim the AI made about a candidate. "Candidate showed strong system design thinking" without supporting quotes was dismissed; the same claim with three transcript quotes was accepted.
- Recruiters were comfortable rejecting candidates based on the AI's report. They were not comfortable advancing candidates based purely on the AI's report — the advancement decision wanted a human-listened audio sample.
That last finding shaped a critical product choice: we made the audio sample player prominent in the report UI, knowing it would be used most for top-of-stack candidates.
What candidates told us — the four findings that mattered
I'm only listing findings where >60% of the qualitative sample independently raised the theme. Lower-frequency findings are in the research appendix.
Finding 1: "I want to know what's happening behind the curtain." The single strongest theme. Candidates who completed the interview reported significantly higher CSAT when they had been told, before the interview, what the AI would and would not do — and what would happen with their data. Candidates who dropped off mid-interview most often cited "I didn't know if this was even being looked at by a human."
Finding 2: "Voice is less awful than I expected. Video is worse than I feared." We had assumed video would be the higher-trust modality (more "human"). The opposite was true. Video introduced anxiety about appearance, lighting, background, and — critically — fear of facial analysis. Voice was rated "professional" and "more like a phone screen." This was a strong-enough signal that we cut video from V1.
Finding 3: "I will lie to a human. I will not lie to a transcript." Several candidates volunteered, unprompted, that knowing the conversation was being transcribed verbatim made them more careful and more honest. This was unexpected. It suggests transcripts function as a soft accountability mechanism in both directions.
Finding 4: "The follow-up questions are how I know it's listening." The single most-cited factor in candidates feeling the interview was "real" was when the AI asked a follow-up question that referenced something specific they had just said. Generic follow-ups ("can you tell me more?") felt scripted. Specific follow-ups ("you mentioned you used Redis for the rate-limiter — what was the eviction policy?") felt like listening. This finding directly drove our investment in adaptive follow-up generation in the LangGraph orchestration.
Affinity mapping — what didn't make the personas
The synthesis surfaced two themes that I deliberately didn't turn into personas, because they are cross-cutting concerns rather than user segments:
- Anxiety about non-native English performance was raised by 9 of 24 interviewees, including some native speakers under stress. This shaped the decision to allow candidates to listen to their own answer back before submitting (in the audio-only mode) and to permit one re-record per question.
- Fear of "AI grading my voice" — concerns about accent, tone, vocal fry, hesitation — was raised by 14 of 24. We do not score audio paralinguistic features. We make this disclosure explicit in the pre-interview flow.
Page 8 — Personas
Four personas, each derived from a clearly differentiated cluster in the research. They are not generic e-commerce templates. Each one has a specific anxiety and a specific opportunity in our product. I'm including persona-specific design implications rather than the usual goals/needs/frustrations format, because design implications are what a persona is for.
Persona 1 — Aarav, the Anxious Junior
Profile: Software engineer, 1.5 years experience, currently employed at a mid-sized Indian SaaS company, applying for senior-junior roles in Bangalore and remote-EU. First time interviewing in 18 months. Confidence: low.
The specific anxiety: Aarav's resume keyword-matches well, but he genuinely doesn't know how he stacks up against the field. He's not sure if he's underselling himself or overselling himself. The AI interview is, for him, the first signal he'll get about his actual market position.
Design implications:
- Pre-interview "what to expect" walkthrough is not optional. Aarav's drop-off risk before starting is the highest of any persona.
- A structured warm-up question ("tell me about a project you're proud of") before any judgment-loaded question lets him calibrate.
- Post-interview, a soft summary of areas he handled well and areas he could prepare more for in future interviews — not a score — converts an evaluative experience into a developmental one. (Decision-quality aside: this is also defensible under DPDP/GDPR's transparency obligations.)
Persona 2 — Priya, the Experienced Skeptic
Profile: Product manager, 7 years experience, currently employed, applying selectively. Has done HireVue once, in 2022, and "never again." Has a low opinion of AI interviews as a category and is doing this one only because the role is genuinely interesting.
The specific anxiety: Priya is afraid of being judged by a machine on dimensions a machine cannot fairly evaluate (judgment, leadership, the taste in a product decision). She is also afraid of looking foolish for taking the interview seriously if it turns out to be a low-effort vendor screening.
Design implications:
- The interview must signal substance within the first 90 seconds — the first question must be open-ended and require real thinking. If Priya hears a generic "tell me about a time when…" she will mentally check out.
- Visible disclosure that the AI does not score paralinguistic features (tone, accent, hesitation). Priya specifically cited this anxiety unprompted.
- A "human review available on request" link in the report-to-recruiter handoff. Priya is the persona most likely to escalate.
Persona 3 — Marcus, the Speed-First Senior
Profile: Engineering director, 12 years experience, currently employed at a US tech company, in active conversations with three competing employers. Time is the binding constraint. AI interviews vs. human screen makes no philosophical difference to him; what matters is the wall-clock cost.
The specific anxiety: Wasting 25 minutes on a screen for a role that turns out to be junior to his level, or for a company that ghosts him after.
Design implications:
- Estimated time prominent in the invitation email, and accurate (median ≤22 min target).
- Allow scheduling at any hour; Marcus does the interview at 10pm after his kids sleep.
- Post-interview, a clear next-step SLA visible to him ("you'll hear from a recruiter within 72 hours, regardless of outcome") converts a low-trust handoff into a contract.
Persona 4 — Eleni, the Visa-Constrained International
Profile: Software engineer, 5 years experience, Greek national currently working in Berlin, applying for roles in Germany, Netherlands, and the UK. Has a complex visa situation. English is her third language and she is acutely aware of accent bias.
The specific anxiety: That the AI is, even unintentionally, scoring her for fluency rather than substance. That the recruiter will see "low confidence" or "hesitant speech" in the AI report and infer something about her ability.
Design implications:
- Explicit disclosure: "we transcribe your answers and evaluate the content, not the audio quality, accent, or speaking style." Repeat this disclosure on the consent screen and in the report.
- Re-record allowed once per question; transcript shown to candidate before final submission for the question.
- For the recruiter: the report flags any claim about the candidate's content with a quoted transcript excerpt. Claims without transcript evidence are not displayed. This matters disproportionately for Eleni because it forecloses the kind of "vibes-based" recruiter pattern-match that disadvantages her.
Why no fifth persona
I considered a fifth — "the underemployed mid-career returner," based on three interviewees with 2+ year career gaps. The cluster wasn't tight enough to be a useful persona; the design implications collapse into Aarav (anxiety) and Eleni (fear of being judged on the wrong dimension). I cut it rather than dilute the four real ones. Cutting personas is a discipline; the temptation in case studies is to keep adding.
Journey map — Eleni, full flow
I'm building this for Eleni because she's the most complex persona and the journey for her tests the design hardest. Aarav, Priya, and Marcus journeys are documented separately.
| Stage | Eleni's experience | Emotion | Design opportunity / what we built |
|---|---|---|---|
| Receives invitation | Email from a Berlin SaaS company. Subject line says "Next step in your application: structured 20-minute interview." Opens, sees AI disclosed in line 1. | Mild relief — disclosure up front. | "AI + Human" framing in subject, plain-language disclosure, time estimate, deadline window (not deadline). |
| Pre-interview prep | Clicks through. Reads what will happen. Sees her data rights clearly: where it's stored (AWS Frankfurt), how long retained (90 days), how to delete. Watches a 90-second sample-question video. | Trust building. | GDPR-aligned consent screen; sample question gives her a chance to calibrate language. |
| Starts the interview | Voice-only. AI greets her, names itself ("I'm Elegant"), confirms consent again, starts with an open warm-up. | Anxiety drops; the first question is real. | Voice-only by design (cut video). Substantive opening question (Priya implication). Live transcript visible to her in a side panel. |
| Mid-interview | AI asks a follow-up that quotes back something she said three turns ago. Eleni notices. | "It's actually listening." | Adaptive follow-up generation via LangGraph. The single highest-impact UX moment. |
| A question goes badly | Her answer to a system design question came out garbled. She uses the "re-record once" option. | Relief, no shame. | Re-record permitted, once per question. Both takes are shown to the recruiter; the second is marked as the candidate's preferred. |
| Completion | Sees a soft summary of topics covered, a confirmation that her recruiter will follow up within 72 hours, and her data rights links again. | Closure. | 72-hour SLA explicit; data rights persistent. |
| Post-interview | 36 hours later, a human recruiter from the company emails her with a follow-up call invite. Recruiter references something specific from her interview. | Validated — a human read this. | Recruiter dashboard surfaces a single "memorable quote" in the candidate card to make follow-up emails specific. |
The stages I am most confident about: the consent flow and the follow-up question moment. The stages I am least confident about: the post-interview summary (we may have shipped this too soft; some candidates wanted a more concrete signal) and the re-record limit of one (some interviewees argued for more, some said even one was excessive).
Page 9 — Three Hard Design Decisions
I want to walk through three decisions in detail, because the value of a case study is in the decisions and their alternatives, not the artifacts produced. For each: the alternatives we considered, what we cut, and the reasoning.
Decision 1 — Conversational vs. structured: we chose hybrid
The choice. Each interview consists of a fixed structured spine (the same 4–6 anchor questions for every candidate for a given role) plus adaptive follow-ups generated dynamically based on the candidate's answer. The structured spine guarantees comparability; the adaptive follow-ups produce signal.
Alternative A we cut: pure structured (all candidates get identical questions, no adaptation). This is what most enterprise screening tools do. It maximizes comparability and minimizes legal risk. We rejected it because Finding 4 from the research was unambiguous: candidates judge the realness of the interview by whether the follow-ups feel real. A purely structured interview signals to candidates that they're being processed, not interviewed, and they disengage.
Alternative B we cut: fully adaptive (each interview is custom). This is closer to how a great human recruiter operates, but it makes comparability across candidates almost impossible, which destroys the recruiter-side value proposition. It also creates a nightmare for bias auditing — you cannot audit a process that varies per candidate. Failed both the recruiter test and the regulatory test.
The hybrid resolves both. Structured spine satisfies recruiter comparability and audit defensibility; adaptive follow-ups satisfy candidate experience. The cost is engineering complexity in the LangGraph orchestration — the follow-up generator has to respect a constraint envelope (don't go more than 2 follow-ups deep, don't drift off the role's competency map, fall back to the next anchor question if the candidate signals confusion).
Decision 2 — Real-time scoring visible vs. hidden: we chose hidden, with a soft post-completion summary
The choice. Candidates do not see any score, signal, or evaluative feedback during the interview. After completion, they see a brief, non-evaluative summary of topics covered and a confirmation of next-step SLA. They do not see how the AI rated them.
Alternative A we cut: visible real-time scoring. Some experimental products (and most coding-interview tools) show candidates how they're doing in real time. We cut this because (a) it amplifies anxiety in Aarav-type candidates to a degree the research showed was disproportionate, (b) it changes the candidate's behavior mid-interview in ways that contaminate the signal, and (c) it almost certainly violates the spirit of GDPR's automated-decisioning provisions when the score is generated by the AI itself.
Alternative B we cut: full evaluative summary post-completion. Candidates would have liked this — Aarav specifically wanted to know how he did. We rejected it because it shifts the legal posture of the report from "evidence for a human recruiter to consider" to "an automated decision that has been delivered to the candidate." That second framing is exactly the one EU AI Act and GDPR Article 22 are designed to constrain. It is also the framing that creates the most reputational risk: imagine the Twitter thread from a candidate who got a 4.2/10 from the AI and then got hired by the same company after the recruiter overrode the score.
The soft summary ("you discussed system design, distributed systems, and team conflict resolution in this interview") is informational without being evaluative. It satisfies Aarav's psychological need for closure without crossing the line into automated decisioning.
Decision 3 — Voice vs. text vs. video: we chose voice-first, text-fallback, no video in V1
The choice. Default modality is voice. Candidates with accessibility needs or unstable internet can switch to text. Video is not offered in V1.
Alternative A we cut: video as default. Higher-fidelity signal, more "human" feel. Cut because the research showed the opposite of our assumption — candidates were more anxious in video, not less, primarily because of fear of facial analysis. Adding video would have compromised the trust thesis of the entire product.
Alternative B we cut: text as default. Lower friction, more accessible globally, no audio-quality issues. Cut because the recruiter research was unambiguous: recruiters wanted to hear the candidate before advancing them. A text-only screen produces a transcript that the recruiter then has to listen to anyway in the next round, which destroys the time-saving value.
The voice-first / text-fallback combination satisfies the recruiter's need for audio (they get an audio sample of every candidate) and the candidate's need for accessibility (text fallback is one click away with no judgment penalty). Video is on the V2 roadmap if and only if we can demonstrate to ourselves that we can offer it without the trust cost.
What we cut from the V1 backlog entirely
A list of features we evaluated and decided against shipping in V1 — included because what you cut tells more about your judgment than what you ship:
- Live coding environment in the interview. Beautiful product surface, real value for engineering roles. Cut because shipping it would have delayed V1 by 4+ months and the use case is narrow. On the V2 roadmap.
- Multi-language interviews. We pilot-tested an English-only V1, fully aware that this constrains us. Localization to Hindi, Bahasa, and German is on the V2 roadmap. The decision was that getting the trust thesis right in one language was more important than getting it half-right in five.
- Recruiter-customizable AI personality. A frequently requested feature. Cut because the moment we let recruiters pick the AI's tone ("be a friendly interviewer," "be a tough interviewer"), we have created a vector for systematic bias that is impossible to audit. The AI's tone is fixed and consistent across customers in V1.
- Candidate-side analytics ("how you compared to other candidates for this role"). Cut for the same reason as Decision 2 — it reframes the report as an automated decision in the candidate's eyes.
- Async candidate Q&A about the role. We tested this. Candidates loved the idea; in practice the AI's answers about the company were too generic to be useful and risked making promises the company couldn't keep. Cut.
Page 10 — Prototype and Usability Testing
Prototyping approach
I'm not going to walk through every Figma frame; the artifacts are linked. What's worth documenting is the iteration sequence, because each iteration was driven by a specific finding rather than by aesthetic preference.
Iteration 1 (paper sketches, 2 days). Two complete flows: a maximalist (every disclosure, every option, every reassurance visible at once) and a minimalist (only the next action visible). Tested informally with three candidates. Both were rated as "too much." The maximalist felt like a legal document; the minimalist felt like a phishing attempt.
Iteration 2 (lo-fi Figma, 1 week). Layered the disclosure into a three-step pre-interview flow rather than one screen. Tested with five candidates in moderated sessions. Completion rate from "click invitation" to "start interview": 100% in this small sample, which made me suspect the test environment was too sterile.
Iteration 3 (hi-fi Figma + real audio simulation, 2 weeks). Wired up a Wizard-of-Oz prototype where the "AI interviewer" was actually a researcher with a voice modulator following a script. This is the iteration where we discovered Finding 4 — the follow-up question moment — was the single biggest trust event.
Iteration 4 (real backend, 4 weeks). First version actually running on the LangGraph orchestration with a smaller-context model behind it. This is when we found that adaptive follow-ups in production were noticeably worse than the Wizard-of-Oz version, which forced us to invest heavily in the follow-up generator's prompt and in the eval harness around it (DeepEval + Ragas + DSPy).
Usability test design
We tested the Iteration 4 prototype with 22 participants (target was 20; over-recruited by 2 to compensate for two no-shows that didn't materialize). Sessions were 60 minutes, moderated, with a think-aloud protocol. Tasks:
- Receive invitation, complete pre-interview disclosure, start interview.
- Complete a 20-minute structured-plus-adaptive interview for a fictional mid-level PM role.
- (Optional) use the re-record feature on at least one question.
- Complete the post-interview summary and indicate next-step expectations.
We instrumented for: time-on-task per stage, drop-off points, errors (any backtracking or asking the moderator for help), think-aloud sentiment markers, and post-test SUS, CSAT, and an open-ended "what would make you not recommend this experience to a friend in your role."
Honest results
Targets vs. actuals. I'm including this table at the level of honesty the deck demands — meaning the metrics that missed are the ones I want to talk about most.
| Metric | Target | Actual | Hit? |
|---|---|---|---|
| Task completion (start → submit) | ≥ 90% | 95% (21 of 22) | ✅ Beat |
| Median time on task | ≤ 22 min | 24.5 min | ❌ Missed by 2.5 min |
| SUS score | ≥ 68 | 71 | ✅ Beat |
| CSAT (1–5) | ≥ 3.8 | 3.9 | ✅ Just hit |
| Errors per task | ≤ 2 | 1.4 | ✅ Beat |
| Drop-off at consent screen | ≤ 5% | 0% | ✅ Beat |
| Drop-off at first question | ≤ 5% | 9% (2 of 22) | ❌ Missed |
What the misses tell me
The time miss (24.5 min vs. 22 target). Driven primarily by two-question stretches where the adaptive follow-up engine asked one too many follow-ups before moving on. This is fixable with a tighter constraint on follow-up depth. A naive fix would be a hard cap at 1 follow-up; the better fix is an LLM-as-judge classifier that decides whether the candidate has produced a complete answer, which is what we shipped in V1.1.
The first-question drop-off miss (9% vs. 5% target). Two participants started the interview, heard the first substantive question, and stopped. Both, in debrief, said the question felt like a "test trap" — they didn't trust their own answer and decided to back out rather than submit a bad one. The fix is not to make the first question easier (Priya's persona warns against that); the fix is to make the first question lower-stakes while still being substantive. The V1.1 change moves the warm-up question's wording from "tell me about a project you're proud of" to "tell me about something you've worked on recently that you found interesting" — softer affective framing, same depth of answer. Drop-off in subsequent testing fell to 4%.
What I'm uncertain about. The CSAT of 3.9 is just over the target of 3.8. With n=22, the confidence interval on that number is approximately ±0.4. Whether we'd see 3.9 or 3.5 in a larger sample is genuinely uncertain. I would not, for example, report this as "we beat our CSAT target" in a board deck without that caveat. The right framing is "candidate satisfaction is in the acceptable range, with confidence-interval-level uncertainty about whether we are at or above target."
Page 11 — Phase I Launch and the First 90 Days
Launch context
We shipped the candidate experience to three pilot customers in late Q1 2025. All three were B2B SaaS companies in the 80–300 employee range, hiring for engineering and product roles. The smallest was running ~25 interviews/month, the largest ~140.
I want to write the next section as carefully as possible, because the easy thing to do here is to cherry-pick the metrics that look good. The honest version is that the launch went better than I feared and worse than the prototype testing predicted.
The numbers
Across the three pilot customers, in the first 90 days post-launch:
| Metric | Pilot result | Target | Verdict |
|---|---|---|---|
| Candidates invited | 1,247 | — | Volume was 12% below pilot capacity sold |
| Completion rate (started → completed) | 71% | ≥ 78% | ❌ Missed |
| Median time-to-completion | 23 min | ≤ 22 min | ❌ Missed by 1 min |
| Candidate CSAT | 3.7 | ≥ 3.8 | ❌ Just under |
| Candidate NPS | +4 | ≥ 0 | ✅ Net positive |
| Recruiter-Action Rate (North Star) | 58% | ≥ 65% | ❌ Missed |
| Latency (p50, candidate-perceived) | Reduced 40% from pilot baseline | — | ✅ |
| Demographic completion gap (largest minus smallest demographic group) | 6.4 pp | ≤ 5 pp | ❌ Missed by 1.4 pp |
The North Star and the equity metric both missed. That matters more than the metrics that hit.
Why North Star missed
The 58% Recruiter-Action Rate was driven by one specific recruiter behavior we did not anticipate: recruiters at the largest pilot customer queued up batches of completed interviews and reviewed them weekly rather than within 72 hours. The product worked as designed; the workflow assumption did not match the reality of how this recruiter operated.
Two responses, both shipped in V1.1:
- We added a recruiter-side aging indicator on completed interviews that visually escalates after 48 hours. This is a soft-pressure design, not a workflow change.
- We added a simple summary email to the hiring manager — not the recruiter — at 72 hours flagging that interviews were waiting. This worked because hiring managers, unlike recruiters, are penalized when their reqs sit. Recruiters cared because hiring managers cared.
After V1.1 the rate moved to 67% over the subsequent 60 days. The lesson is that North Star metrics should be designed against actual user workflows, not idealized ones.
Why the equity metric missed
The 6.4 percentage point completion gap was between two groups: candidates whose primary language was English (highest completion) and candidates for whom English was a third or later language (lowest completion). The gap was driven by drop-off concentrated in the first technical question rather than throughout the interview.
Investigation suggested the technical question was using English idioms ("set this up from scratch," "walk me through your thinking") that translated as more cognitively loaded for non-native speakers than a literal phrasing would. We rewrote question stems to use literal-translation-friendly phrasing in V1.1. The gap closed to 3.8 pp in subsequent testing.
I want to be specific about the limit of this fix: closing a 6.4 pp gap to 3.8 pp does not mean we have eliminated the bias. It means we have reduced one observable channel of it. The deeper question — whether the AI's evaluation of substance is itself language-fluency-correlated — is open. We measure it; we cannot yet claim to have solved it.
What surprised us
Positive surprise: The transcript-as-artifact decision (Decision 1 in the framing of the case study, originally a Decision 2 in product) turned out to be more commercially valuable than we expected. Two of the three pilot customers said in their renewal conversations that the transcripts were the single most valuable feature, because they enabled hiring managers to review candidates without scheduling a synchronous review. We had under-marketed this and had to rewrite the demo flow.
Negative surprise: Candidates who chose the text fallback (vs. voice) had a 14 percentage point lower completion rate. We expected text to be a low-friction accessibility option; in practice, it was being chosen by candidates who didn't trust the voice modality, and that low trust manifested in completion. The text fallback now defaults to a more guided, confidence-building flow with explicit reassurances. Completion in the text path is still lower than voice but closing.
Page 12 — Reflection and Forward Thinking
What I'd do differently
I would have started with the recruiter workflow research, not the candidate research. The North Star miss was not a product problem; it was a workflow assumption problem, and we would have caught it in week 2 of recruiter shadowing, not week 12 of pilot launch. The candidate research was strong but I over-prioritized it.
I would have tested the equity metric earlier. The 6.4 pp completion gap on language fluency was visible in the prototype testing; we had the signal. We did not treat it as a launch-blocking issue at the time, and we should have. It's the metric I think about most.
I would have shipped V1 with explicit language preference selection. Letting candidates choose between "interview me in English" and "interview me in English with literal phrasing" (or, eventually, in their native language) is a feature we deferred. In retrospect it should have been V1, not V2.
Second-order thinking — what happens next
Candidates will start preparing for AI interviews specifically. This is already happening. Tools like Final Round AI and Ultracode coach candidates through AI interviews, sometimes by literally answering for them in real time. Our response in V2 is not a coaching-tool detector (an arms race we'd lose) but a structural one: questions that are difficult to answer with prepared scripts because they probe specific transcript content from earlier in the same interview. The follow-up engine becomes the moat.
Recruiters will trust the AI report progressively more, and their independent screening muscles will atrophy. This is a longer-horizon risk. A recruiter who runs 200 candidates a quarter through Elegant V1 will, in 18 months, be a worse independent screener than they were in month 1. We don't have a clean answer for this. We are experimenting with a "spot-check" mode where the AI report is occasionally withheld for ~5% of candidates, forcing the recruiter to do an unaided review and producing an inter-rater reliability metric over time. Whether this will survive customer pressure to "show me the report every time" is an open question.
Per-interview pricing changes recruiter behavior in ways we have to watch. When you charge per interview, customers economize on interviews. The healthiest version of this is "they only invite genuinely qualified candidates." The less healthy version is "they avoid using the product on borderline candidates and just reject them," which removes exactly the candidates the product is most useful for. We are tracking the ratio of invited-to-applied as a leading indicator of this drift.
The "AI + Human" framing has a sell-by date. Right now it is a defensible regulatory and reputational stance. In 24–36 months, as AI hiring decisions become normalized in some jurisdictions and more strictly regulated in others, the framing will need to evolve. The sustainable version is not "AI + Human" as a permanent positioning but as a transitional one toward something more nuanced — probably "AI for evidence, human for decisions" with explicit tooling for which decisions a human is the only acceptable maker of. That distinction is where the next year of product work lives.
Third-order thinking — what this becomes in 5–10 years
I want to be careful here, because most third-order thinking in case studies is grandiose. I'm going to keep it to three claims I genuinely believe.
(1) The transcript dataset becomes the strategic asset, and the regulatory liability, simultaneously. Every candidate who completes an Elegant V1 interview produces a transcript that contains structured, real-time, comparable evidence about a labor market. Aggregated across customers and time, this dataset is one of the most detailed labor-market signals ever collected — it is what the BLS would build if the BLS could conduct interviews. The commercial possibilities (predictive models for compensation benchmarks, skill-trajectory mapping, internal mobility recommendations) are large. So is the regulatory exposure. Under DPDP and GDPR the secondary use of this data is constrained; under EU AI Act provisions, the model trained on this data inherits high-risk classifications. The asset and the liability are the same data. The work to separate them — through privacy-preserving aggregation, federated approaches, customer-by-customer data isolation — is multi-year and is more important than the next product feature.
(2) AI-mediated hiring will erode the apprenticeship model for junior recruiters. This is the pattern I worry about most, because it is invisible to our customers in real time. When the AI conducts the screening interview, the junior recruiter who would have learned screening by doing it does not learn it. We are probably in the first generation of a workforce of senior recruiters who never did the foundational work that built their senior counterparts' judgment. The same dynamic is unfolding in junior software engineers and AI assistants. Whether this is a real loss or a generational adjustment is not yet knowable. As a product company we benefit from the displacement; as a society I think we are running an uncontrolled experiment.
(3) The marketplace model will eventually compete with the SaaS model on the same axis. Right now Elegant V1 (SaaS) and Mercor (marketplace) sell to different customers — we sell to teams that have a funnel and need to process it, they sell to teams that don't have a funnel and need one. In 3–5 years these converge. Mercor adds tooling for customer-owned funnels; we add candidate sourcing. The eventual competitive question is whether the labor-market data moat (us) is more durable than the candidate-supply moat (them). I think it is, but I hold the view loosely.
Team
Phase I shipped because of:
- Shivam Gupta (Co-Founder), who built the orchestration backbone (LangGraph, n8n, AWS migration) and held the eval-harness discipline that made the adaptive follow-up generator work in production. The 40% latency reduction is his work as much as anyone's.
- The engineering team, who shipped the candidate-facing flow against an aggressive timeline.
- Our three pilot customers, who tolerated the pilot rate of breakage and gave us the recruiter workflow research that the product needed.
- Every candidate who completed a pilot interview and then sat for a debrief with us. That debrief data is the foundation of every page above this one.
I also want to be honest that I made mistakes I haven't documented here in detail. The recruiter-workflow miss was the largest of them. The equity-gap delay was the most consequential. Phase II's research plan is structured to catch the analogues of these earlier.
Where this goes next
Phase II scope (in order of priority):
- Recruiter dashboard and workflow features — the under-investment of Phase I.
- Multi-language interviews — Hindi, Bahasa, German.
- Live coding environment for engineering roles.
- Spot-check mode for recruiter calibration and inter-rater reliability over time.
- Privacy-preserving aggregation tooling for the transcript dataset.
The decision about whether to add video to V2 is deferred. The trust thesis worked because we did not ship video. Adding it back is a non-trivial reversal that I am not yet convinced we should make.
Document version: Phase I retrospective, draft for portfolio. — Mayank Lakhani, Co-Founder & Product Lead, RansahAI.