AI / MLPrototype2026
Shortlist
Resume screening where every criterion is a typed question to TypeSafe's Jev: add a new requirement and 40 resumes are re-scored on it in about a second, for a tenth of a cent.
- 80
- Synthetic resumes across 3 openings
- ~1 s
- To score 40 resumes on a new criterion
- 0.94
- Composite fit against planted skill levels (r)
Timeline
Oct 2026
Role
Design, scoring model and front end
Team
Solo
Domain
HR tech · Recruiting
Status
Prototype
Stack
Source
Section 01
Overview
LLM resume parsers extract a fixed set of fields up front. When a hiring manager asks for something new, every resume has to be parsed again, which is slow and expensive.
Shortlist treats each criterion, like "Java and Spring depth" or "Notice period 30 days or less", as its own typed question to Jev, TypeSafe's decision model, with a written rubric. Answers are cached per resume and question, and the ranking is plain arithmetic in the browser, so changing a weight re-ranks instantly with no model call at all.
Section 02
A New Criterion, Now
A recruiter adds a requirement in one click or writes their own as a scale or a yes/no question. Only that question is sent, one request per resume with 24 in flight, and the new column fills in live. In measured runs, adding a criterion to 40 resumes took 0.5 to 1.4 s and cost about $0.0012; a full nine-criterion screen of 40 resumes took 0.3 to 1.8 s for $0.0023.
Section 03
Every Score Explains Itself
Each answer shows its full probability distribution across the rubric levels and a confidence label, and the composite fit breaks down line by line as weight times answer. Candidates can be compared side by side, and nothing is rejected automatically: moving a candidate through the pipeline is always a recruiter's decision.
Section 04
Does It Read Resumes Correctly?
The sample resumes are generated from hidden skill levels, which makes a check possible. Jev's scores tracked those levels closely, with correlations of 0.92 to 0.99 on the skill criteria and 0.94 for the overall fit. The weakest was resume clarity, at 0.73. Because the resumes are built from templates, this shows Jev reads the planted signal well; it is not evidence about real resumes.
Section 05
Before Real Hiring
Screening real people needs safeguards this prototype does not have yet:
- Redact names, contact details, location and other identity fields before any question is asked
- Review the suggested criteria for proxies such as location or employer prestige
- Measure score calibration and disparity on real, consented data
- Keep an audit log of every criterion, weight and decision









