These days, with AI, assessments still continue to test code output. Meanwhile, at each of our companies, though, PR counts have nearly tripled. Most of us haven’t manually edited a line of code in a year.
Our teams are putting more and more emphasis on code and architecture reviews, yet hiring processes haven’t changed whatsoever. Even the AI-assisted ones we’ve done still assess code output as the primary evaluation metric. As agents develop, we truly believe this will not be the most challenging ability for an engineer to have.
We’ve seen that the real difficulty with using AI is not just reviewing code your AI generates; it is reviewing code that another engineer’s AI has generated, having little context yourself.
That’s why we’re launching Merge today. Our goal with Merge is to assess engineering judgement. We're still quite early in the design phase, but here’s how it works:
1. Candidates are shown a small codebase to understand and a PR to review and comment on. 2. An AI agent addresses each PR comment via a code change or reply, simulating a real engineer. 3. Candidates can repeat until 5 revisions are used up or time runs out.
At the end, we assess the following: 1. Coverage - How many bugs or vulnerabilities did the candidate identify and address? 2. Communication - Was the candidate efficient and constructive with their feedback? 3. Efficiency - How many revisions and tokens did the review take?
We’re the first platform that can show companies exactly how efficient a candidate is with token use, LLM costs, and PR revisions — all of which are exceedingly important in real jobs.
If you’re interested in the next-generation of engineering hiring, book a demo with us! Feel free to ask any questions below as well.
kasts•40m ago
harshithl1777•30m ago