frontpage.
newsnewestaskshowjobs

Made with ♥ by @iamnishanth

Open Source @Github

Open in hackernews

Ask HN: What tools are you using for AI evals? Everything feels half-baked

4•fazlerocks•22h ago
We're running LLMs in production for content generation, customer support, and code review assistance. Been trying to build a proper evaluation pipeline for months but every tool we've tested has significant limitations.

What we've evaluated:

- OpenAI's Evals framework: Works well for benchmarking but challenging for custom use cases. Configuration through YAML files can be complex and extending functionality requires diving deep into their codebase. Primarily designed for batch processing rather than real-time monitoring.

- LangSmith: Strong tracing capabilities but eval features feel secondary to their observability focus. Pricing starts at $0.50 per 1k traces after the free tier, which adds up quickly with high volume. UI can be slow with larger datasets.

- Weights & Biases: Powerful platform but designed primarily for traditional ML experiment tracking. Setup is complex and requires significant ML expertise. Our product team struggles to use it effectively.

- Humanloop: Clean interface focused on prompt versioning with basic evaluation capabilities. Limited eval types available and pricing is steep for the feature set.

- Braintrust: Interesting approach to evaluation but feels like an early-stage product. Documentation is sparse and integration options are limited.

What we actually need: - Real-time eval monitoring (not just batch) - Custom eval functions that don't require PhD-level setup - Human-in-the-loop workflows for subjective tasks - Cost tracking per model/prompt - Integration with our existing observability stack - Something our product team can actually use

Current solution:

Custom scripts + monitoring dashboards for basic metrics. Weekly manual reviews in spreadsheets. It works but doesn't scale and we miss edge cases.

Has anyone found tools that handle production LLM evaluation well? Are we expecting too much or is the tooling genuinely immature? Especially interested in hearing from teams without dedicated ML engineers.

Comments

PaulHoule•22h ago
I worked at more than one startup that was trying to develop and commercialize foundation models before the technology was ready. We didn't have the "chatbot" paradigm and were always focused on evaluation for a specific task.

I built a model trainer with eval capabilities that I felt was a failure, I mean it worked, but it felt like a terrible bodge just like the tools you're talking about. Part of it is that some the models we were training were small and could be run inside scikit-learn's model selection tools which I've come to seen as "basically adequate" for classical ML but other models might take a few days to train on a big machine which required us to develop inferior model selection tools that worked with processes too big to fit in a single address space but also gave us inferior model selection for small models. (The facilities for model selection in hugginface are just atrocious in my mind)

I see a lot of bad frameworks for LLMs that make the same mistakes I was making back then but I'm not sure what the answer is, although I think it can be solved for particular domains. For instance, I have a design for a text classifier trainer which I think could handle a wide range of problems where the training set is between 50-500,000 examples.

I saw a lot of lost opportunities in the 2010s where people could have built a workable A.I. application if they were willing to build training and eval sets and they wouldn't. I got pretty depressed when I talked to tens of vendors in the full text search space and didn't find any that were using systematic evaluation to improve their relevance. I am really hopeful today that evaluation is a growing part of the conversation.

VladVladikoff•4h ago
>We're running LLMs in production for content generation, customer support, and code review assistance.

Sounds like a nightmare. How do you deal with the nondeterministic behaviour of the LLMs when trying to debug why they did something wrong?

Ask HN: Any good tools for viewing congressional bills?

17•tlhunter•51m ago•9 comments

Ask HN: What are some good resources for coding best practices?

3•genericmask•1h ago•1 comments

Ask HN: Should I build a directory product?

3•alizaid•1h ago•2 comments

Ask HN: Startup getting spammed with PayPal disputes, what should we do?

275•june3739•2d ago•179 comments

Ask HN: Has anybody built search on top of Anna's Archive?

282•neonate•2d ago•146 comments

Ask HN: Anyone else feeling increasingly alienated from the industry?

15•saubeidl•8h ago•14 comments

Ask HN: What are your fav/goto decision making hacks/heuristics?

4•ottaborra•8h ago•7 comments

Ask HN: Who is hiring? (June 2025)

365•whoishiring•4d ago•457 comments

Ask HN: Running AI agents in isolated environments

4•polycaster•10h ago•0 comments

Ask HN: How do I learn robotics in 2025?

395•srijansriv•4d ago•99 comments

Ask HN: How do I learn practical electronic repair?

180•juanse•6d ago•112 comments

Ask HN: Anyone making a living from a paid API?

247•meander_water•6d ago•172 comments

Ask HN: What do you put in claude.md and what you leave out?

5•bognition•1d ago•2 comments

Ask HN: Options for One-Handed Typing

92•Townley•2d ago•93 comments

Ask HN: Walking while working and having meetings

3•martythemaniak•14h ago•5 comments

Ask HN: Who's Using the Origin Private File System?

4•ChadNauseam•15h ago•1 comments

Ask HN: Who wants to be hired? (June 2025)

125•whoishiring•4d ago•384 comments

Ask HN: What Does Your Self-Hosted LLM Stack Look Like in 2025?

17•anditherobot•1d ago•6 comments

O(1) memory, no-preprocessing reachability algorithm for 2D grids

2•MatthiasGibis•21h ago•1 comments

Ask HN: What tools are you using for AI evals? Everything feels half-baked

4•fazlerocks•22h ago•2 comments

Ask HN: Where do you go for cutting-edge dev news and info?

2•TimTheTinker•1d ago•9 comments

Ask HN: What is the best LLM for consumer grade hardware?

238•VladVladikoff•1w ago•182 comments

Ask HN: How are parents who program teaching their kids today?

101•laze00•5d ago•91 comments

Ask HN: Dealing with Vibe Coding Depression?

16•softirq•1d ago•23 comments

Reaching my first 100 users without money or audience (at 10K users now)

32•felixheikka•3d ago•11 comments

How do you store and maintain your CV/resume over time?

10•xantin•2d ago•16 comments

Ask HN: List of skills to survive the AI tsunami

16•cookiemonsieur•1d ago•6 comments

Ask HN: What's with the repeated job posts on "Who's hiring"?

85•rafavento•2d ago•41 comments

Ask HN: Is Adrian Colyer of "The Morning Paper" fame ok?

10•yencabulator•16h ago•0 comments

Ask HN: Best way to get laid off

10•jakamm•1d ago•24 comments