frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

ICANN Reveals 2026 Round Applications for New Generic Top-Level Domains

https://www.icann.org/en/announcements/details/icann-reveals-2026-round-applications-for-new-gene...
1•ChrisArchitect•42s ago•0 comments

Panolayer – Agent code comprehension and deterministic detection of regressions

https://panolayer.com
1•jederle•4m ago•1 comments

gTLD Reveal Day: 1372 unique strings

https://newgtldprogram-aps.icann.org/applications
1•bwblabs•6m ago•2 comments

Show HN: Smart Blur – Auto-Blur PII in the Browser, with a OpenAI Local Model

https://smartbuildlabs.com/apps/smart-blur/
2•vickyonlinecont•7m ago•2 comments

Claude: Monthly API credits for Max and Team plans

https://support.claude.com/en/articles/17154008-monthly-api-credits-for-max-and-team-plans
3•stigi•10m ago•0 comments

Cloud Maturity Is the New Digital Transformation

https://phpscientist.com/blog/cloud-maturity-digital-transformation-ai-ready-enterprise/
1•senthil_kr•10m ago•0 comments

Muere la puta vieja los cojones

https://www.elmundo.es/madrid/2026/10/07/6ac68ee8fc6c8323468b4599.html
1•UltimisimaHora•11m ago•0 comments

The world has nearly burned through its oil stockpile buffer

https://www.reuters.com/business/energy/world-has-nearly-burned-through-its-oil-stockpile-buffer-...
3•geox•11m ago•0 comments

Skate 3 in the Browser

https://skate.aaddpp.lol/
1•E-Reverance•12m ago•0 comments

Meta and Microsoft Limit Employee Use of Claude AI Tools

https://www.rswebsols.com/news/meta-and-microsoft-take-steps-to-reduce-employee-usage-of-claude-ai/
2•speckx•12m ago•0 comments

Strands Box: The Big Picture

https://strandsagents.com/blog/strands-box-the-big-picture/
1•saikatsg•14m ago•0 comments

Building Windows for Hybrid Intelligence

https://blogs.windows.com/windowsexperience/2026/10/07/building-windows-for-hybrid-intelligence/
1•antimora•15m ago•0 comments

The Barriers of Perception

https://terrytao.wordpress.com/2026/10/07/the-barriers-of-perception/
1•smilelamp•15m ago•0 comments

Get your AI agent connector or integration discovered on TryMuse

https://trymuse.com/
1•Tonje•16m ago•0 comments

New gTLD Program: 2026 Round Application Statistics

https://newgtldprogram-aps.icann.org/statistics
1•jaredwiener•18m ago•0 comments

Push Ifs Up and Fors Down: The Idiom, Its Algebra, and Its Limits

https://debasishg.github.io/blog/push-ifs-up-fors-down/
2•speckx•19m ago•0 comments

Ask HN: What will be the next AI wave?

1•mathieu_aithos•20m ago•0 comments

OpenDocRouter: A Unified API for OCR Models

https://www.opendocrouter.ai/
1•cheesyFish•21m ago•0 comments

Everyone Is an Engineer Now

https://www.jacob.co/blog/everyone-is-an-engineer
4•jacobsimon•21m ago•0 comments

Show HN: TraceUX – Self-hosted session replay in one Go binary

https://github.com/fjosue4/trace-ux
3•fjosue4•21m ago•0 comments

Proteus: A hypervisor-agnostic VPC for Triton Cloud

https://tritoncloud.io/blog/proteus-a-hypervisor-agnostic-vpc-for-triton-cloud/
2•NexRebular•21m ago•0 comments

Moon – A mood tracker where the app does the math and the LLM only writes

https://moonapp.lv/
4•nikkey80•23m ago•0 comments

A hallucinated module, a backfiring RAG pipeline and the MCP server to fix it

https://tmtabor.io/blog/genepattern-copilot/
3•tmtabor•25m ago•0 comments

Measure how often coding agents choose your devtool

https://agentpicks.io
5•iacguy•25m ago•0 comments

Ask HN: Will mental health improve as AI brings mathematicians down to earth?

3•amichail•26m ago•0 comments

'Artificial' Took on a Big Tech Giant. Hollywood Wasn't Ready

https://www.hollywoodreporter.com/movies/movie-features/artificial-exclusive-luca-guadagnino-andr...
2•laurex•26m ago•0 comments

What Will Happen to Indian IT? Past and Future of the "world's back office."

https://asteriskmag.com/issues/15/what-will-happen-to-indian-it
3•pseudolus•27m ago•0 comments

I indexed 109k events across 50 cities – 83% don't publish a price

https://arclight.events/data
2•Siddhant8019•28m ago•0 comments

Show HN: Wrapper for Free Open Router Models

https://wfform.com/
1•metaph6•29m ago•0 comments

Reverse Easter Egg from 2012

https://tobeornottobe.com
1•sans_souse•29m ago•0 comments
Open in hackernews

Claude Haiku 5.5

https://www.anthropic.com/claude-haiku-5-5
236•sfkgtbor•1h ago

Comments

TheAmazingRace•57m ago
I wonder if we have an AI LLM equivalent to Moore's Law. Like how often do we expect improvement in this technology and with what timing?
himata4113•57m ago
double the information density every 2 days?

serious bit: if you think about how these smaller models work, at the end of the day it seems that they are now capable of forgetting useless information because they're able to derive it in reasoning allowing models to become smaller at the cost of requiring more reasoning tokens to solve a task.

qeternity•53m ago
Knowledge will be shifted to systems like n-gram augmentation which are relatively cheap and will not compete with reasoning capabilities for weight saturation.
dyauspitr•55m ago
Hopefully enough runway for an existing model to train the next to be better than itself with absolutely no human intervention.
bravetraveler•54m ago
I've heard tell about 100% of certain types of work being ended in batches of six months. For years. Truthfully, I'm skeptical, but accuracy wasn't prioritized.
ChaseRensberger•54m ago
reminds me of this blog post: https://campedersen.com/singularity
onlyrealcuzzo•44m ago
Yes -> every 18 months they've gotten 90% more efficient for the same level of quality for about 5 years. There's little sign that trend is slowing. If anything, there's reason to believe that System 1 models (plus potentially 1-2-3 workflows) may increase that over the next 3-5 years.

You'll know when the trend stops -> when the intelligence differential between smaller models like 7B starts to grow instead of shrink from 32B models -> that means 7B is getting about as smart as it can get. Then, 32B will follow next, then 70B, etc etc.

We haven't yet seen that at any size AFAIK.

thefourthchime•40m ago
Andrej Karpathy said once that he expects superintelligence could fit in 1 billion parameters.
onlyrealcuzzo•22m ago
Super intelligence that doesn't have to deal with the real world, maybe.

I wouldn't be surprised if less than 1B param equivalent of our brain deals with solving math and writing computer programs and physics and all the things we tend to associate with "intelligence" - especially if you ultra optimized for that, I doubt our brain works like that.

Dealing with the real world, I highly highly doubt it.

jstummbillig•15m ago
How about if we get away from written text as the input, to something more fundamental, that then also is able to produce text (among other things)?

Given that humans learn to talk while having encountered a measly number of word instances, and, given enough time, we should always be able to improve on the lottery that is biology, it does seems fairly likely.

istjohn•36m ago
According to Epoch AI:

> The cost of achieving a given level of AI performance has fallen about 47% per quarter since 2023, or 13× per year. [0]

0. https://epoch.ai/publications/the-plunging-price-of-thought

FooBarWidget•16m ago
Then why are AI plans still so super expensive, and AI spending going through the roof, while all the subsidies are ending?
jstummbillig•14m ago
Because it's increasingly useful and the thing you are displacing (human time) is much more expensive.
teaearlgraycold•6m ago
At least for me the Claude plans seem like an incredible deal and I never hit my limit.
adgjlsfhk1•4m ago
The cost per fixed level of intelligence is dropping, but we're also getting dramatically more intelligent models.
sfkgtbor•57m ago
Happy about the Sonnet cache read price cut.
minimaxir•50m ago
That was effectively required to match GPT-6.1 Sol (costs and caching prices are now equal). Sonnet 5.5 made zero sense to use over Opus 5.5 under the old cache prices.
minimaxir•56m ago
Pricing is...a bit weird.

    Input
    $0.10 / MTok for prompts up to 100,000 tokens
    $0.50 / MTok for prompts over 100,000 tokens

    Output 
    $0.50 / MTok for prompts up to 100,000 tokens
    $2.50 / MTok for prompts over 100,000 tokens
100k tokens is an absurdly low cutoff and it is only applicable to Haiku and not Sonnet or Opus. It's a low enough cutoff that it will be quickly exceeded if you are doing anything with Agents; for typical generation or Jev-like classifiers, it's a good value and as noted in this article, that is apparently the vast majority of Haiku use.

In both cases, still much cheaper than Haiku 4.5's $1 input / $5 output and these prices better compete with GPT-6 Luna. ($0.10 input / $0.50 output, but with no token threshold [EDIT: the threshold for Luna is apparently 272k])

giancarlostoro•53m ago
I with they'd give Haiku like 400k tokens roughly, I think between 400k or even 600k tokens is a sweet spot, but Haiku is basically designed to be for small edits is my understanding, but it sucks because any time I ask Opus to "try" letting Haiku do the work, it just falls apart and Opus comes back and tells me it switched to Sonnet (even before Sonnet finally jumped up to 5.x).

I will try the new Haiku, but it would be worthwhile if Haiku could take sane instructions and do all file editing for Opus / Sonnet / Fable then it would be worth using.

tr4656•52m ago
Luna does as well, but just at a higher limit.

From OpenAI's website: Prompts with more than 272K input tokens are priced at 2x input and cache rates and 1.5x output for the full request.

minimaxir•47m ago
Huh, that disclaimer is on the model page (https://developers.openai.com/api/docs/models/gpt-6-luna) but not the pricing page. Annoying.

Fixed.

maz1b•54m ago
Wow, the rate of improvements in the AI era is staggering.

GDPval-AA v2.1 as of now: 1620

GDPval-AA v2.1 for Haiku 4.5: 735

The 100k tokens pricing makes sense, looks to be a hedge against OpenAI's decisions API and Jev or its open source alternatives that are springing up.

Nice release, congrats to Anthropic.

sroussey•52m ago
It’s about time Haiku got an update!
iagocc•52m ago
Where is Pelican? ehehhe
rvz•47m ago
This is what AI psychosis has reduced readers to on this site.

We are witnessing the acceptance of average and accelerating more of the same low quality slop.

InsideOutSanta•43m ago
So do you have the pelican or no?
swalsh•39m ago
I think it's a joke at this point, but also the visual benchmark is a remarkably dense method for demonstrating how good a model is.
InsideOutSanta•19m ago
Yeah, people like to poop on the pelican. But pelican quality still correlated with overall model capabilities reasonably well, and you can immediately see and interpret it. It's a running gag, but it also does have some actual value.
TomGarden•51m ago
From these selected benchmarks, it looks like it smokes Luna capability-wise. Excited to put it through its paces
margorczynski•51m ago
How does the price compare to Luna? At least looking at the numbers it is noticeably better at most tasks.
onlyrealcuzzo•50m ago
IMO, this is better. Luna is super cheap, but it's not that capable. At higher levels of reasoning, it's not that fast.

This is more expensive, but it also looks like it's better enough that it's far more useful.

I also won't be surprised if you look at cost per completed task + wall clock time that it comes out ahead for the majority of what you'd want to actually use it for.

Luna will still be a great option for doing non-engineering tasks super cheaply.

TomGarden•48m ago
For prompts under 100k tokens, it's priced the same as Luna - $0.10 in, $0.50 out.

For prompts over 100k tokens it's 5 times more expensive - $0.50 in, $2.50 out.

simianwords•50m ago
I remember a friend asking me why LLMs suck so bad. She was using Haiku 4.5 and that poor model couldn't keep track of the context within 3 messages.

She said she was using Haiku 4.5 because she was advised to be careful with the spending.

I hate that model so much lol.

seaal•50m ago
The monthly API credits for Max plan seems fantastic, especially considering Haiku pricing. Being able to actually use my Claude plan for other harnesses and use-cases on top of regular CC usage is everything I wanted.

Anthropic has really been doing all the right things in the past few weeks, while OpenAI continues to fumble the bag.

0gs•44m ago
yeah totally agree. esp how efficient it can be to have a subscription quota-paid orch spin up a bunch of API agents, this is kind of like free money to encourage what was already an easy way to save money (via batch pricing)
thepasch•29m ago
Note that this is Anthropic Trojan-Horsing the previously announced June change in with a model release, where the Claude Agent SDK can no longer be used with Claude subscriptions and is now billed with API credits only.

https://support.claude.com/en/articles/15036540-use-the-clau...

InsideOutSanta•22m ago
Ah, that sucks. I'm using Paseo to run Claude Code; I guess that just got a whole lot more complicated.
patrickwdaly•50m ago
How are y'all using Haiku though? I rarely select it.
svachalek•46m ago
Opus often picks it when it's doing a "find me something" subagent. But largely it's been held back by being fully a year old at this point, and priced at a much higher price than models that are far more capable.
steve_adams_86•46m ago
It's great at parsing documents inexpensively. For the few skills/plugins I've made, I usually instruct Claude to use Haiku for low-reasoning grunt work.
mariocesar•43m ago
I have a zsh functions that calls claude code with haiku to suggest commit messages, is faster and the instructions are two lines.

I also have an "ask" script that I use daily to ask simple stuff, it can access websearch and webfetch, it's more than enough to parse logs, ask for commands, quick research on the internet, small stuff. https://github.com/mariocesar/dotfiles/blob/main/common/.loc...

I use haiku for things that needs to be quick, have really clear instructions.

gghootch•41m ago
I was waiting for this.

Planning on doing flash analyses of PRs that impact evals in some way, and then post comments on GitHub whenever there’s flaws in them

( https://evalship.com )

tpoacher•49m ago
Good to see Anthropic back alternative OSes.
afrnswrth•48m ago
The important question though...how does it do making a pelican on a bicycle?
caaqil•46m ago
> Haiku 5.5’s cybersecurity safeguards are more restrictive than Haiku 4.5’s, but somewhat less restrictive than those we’ve applied to other recent models. In cybersecurity, they permit a wider range of defensive tasks than our safeguards for Sonnet 5.5, but they still block penetration testing and other techniques more likely to be used by attackers.

If you block pentest or "other techniques more likely to be used by attackers", then what does "permit a wider range of defensive tasks" even mean?

Any defensive task that's meaningful is almost indistinguishable from legitimate red-teaming that then falls under 'likely to be used by attackers". If only they would just stop nerfing these models, that'd be great. No APT is waiting around for Anthropic's permission, so might as well let us have some cool stuff.

simianwords•45m ago
> Second, this week, we’ll roll out a new monthly API credit to all Max and Team subscribers for use on the Claude Platform. Max 5x users will get $100 in credits per month, Max 20x users will get $200, and Team subscribers will receive up to $500, pooled across their users. These credits are designed to allow our users to experiment with building tools, apps, and agents that call our API. They can be used on any of our models. For more information, see our Help Center article.

Did anyone read this? We get free API credits on some plans now

swalsh•43m ago
Top of the page in 17 minutes? Now I know what y'all do while your agents are working.
AtNightWeCode•42m ago
Probably the same scam as the last Haiku update I guess. Uses more tokens to compensate for the lower price.
charlesabarnes•40m ago
> Second, this week, we’ll roll out a new monthly API credit to all Max and Team subscribers for use on the Claude Platform. Max 5x users will get $100 in credits per month, Max 20x users will get $200, and Team subscribers will receive up to $500, pooled across their users

This is a very big benefit for me. I can now ship actual ai enhanced features behind my subscription without paying extra or fully relying on on-device models. I do worry that this is to soften the blow for user-unfriendly changes

geek_at•33m ago
This is literally for you to get tangled in their api and when they stop giving you the allowance they hope you will just continue to pay
enraged_camel•23m ago
What does "tangled in their api" mean? Switching is pretty easy.
charcircuit•11m ago
Not really, you have to fiddle with generating api keys and setting environment variables. Meanwhile with Anthropic it will just start charging you API prices for the tokens you are generating without even a single warning.
enraged_camel•5m ago
>> Not really, you have to fiddle with generating api keys and setting environment variables.

That's 5-15 minutes of work at most. Not exactly the type of lock-in the parent is implying.

tomjen3
vickyonlinecont•39m ago
Does anyone still use Haiku model?
jstummbillig•39m ago
Opus 5.5 does :^)
crooked-v•38m ago
The important question is, does it talk in incomprehensible Claude-ese like the other Claude 5.x models?
garo-pro•38m ago
> Claude Haiku 5.5 is our fastest model to date at each model’s standard speed, although it runs less quickly than our Opus models in Fast Mode.

Opus 5.5 runs 117 tps average on Openrouter, so it must be at least 10-20 tps slower for them to mention. IDK why they mention this as it does not help for marketing though. https://openrouter.ai/anthropic/claude-opus-5.5

jstummbillig•37m ago
Maybe they think it's of interest.
wyrdcurt•33m ago
About time Anthropic released a competitive cheap model. Haiku 4.5 has been too expensive compared to its performance for months now (in fact I don't remember being too impressed even when it was released). This one actually looks worth using in some scenarios. If it's really as much of a step up from Luna as the benchmarks they've shown indicate, it'll probably replace Luna in my workflows. 100k tokens is a pretty low threshold before the price goes up, but I tend to use these smaller models for smaller tasks anyway.
jjcm•32m ago
Ran image -> html tests for this. I was curious if this smaller model was good enough for complex UI. It was not.

Haiku 5.5: https://html.non.io/lcars-haiku-5.5/

Opus 5.5 for comparison: https://html.non.io/lcars-opus-5.5

Designs it was building from: https://diffui.ai/app/canvas/5093e689-1e74-4f26-b632-2a4500f...

One interesting thing is it took a look at the job at hand, and immediately delegated it to Opus 5.5. It at least knows what it isn't good at. Very fast though, and likely best used for small subagent tasks / tightly scoped work.

BrokenCogs•27m ago
Neither of these look "good" to me. There is so much visual noise on the page, like someone turned the "AI Slop" dial to 11. In fact I prefer the simpler design Haiku made.
FranzFerdiNaN•18m ago
It’s not really AI slop, it’s how most modern SAAS websites look like.
twostorytower•16m ago
It's not really about whether the design looks good. It's about if the model can take the design given to it and replicate it in code. Opus 5.5 matches the designs almost to the pixel. Haiku built something else entirely.
jjcm
bouk•27m ago
This is great! Been using GPT 6 Luna for decompiling my childhood favorite game (Age of Mythology) and this means I can throw Haiku into the mix as well. 17352/21965 functions matched so far...
gizmodo59•20m ago
can you share more details? was this very involved or asking codex/claude/open code with a 1 shot like approach?
bouk•6m ago
I'll write a blogpost when I actually have it working, but basically I gave the game .msi installer to claude opus 5.5 and said to read these blogs:

  - https://blog.chrislewis.au/using-coding-agents-to-decompile-nintendo-64-games/
  - https://blog.chrislewis.au/the-long-tail-of-llm-assisted-decompilation/
And to setup a harness that will decompile the game and start doing a matching decompilation of every function. It set up a bunch of tooling and started a service in the background to do this actual decompilation campaign. I put some instructions into the main opus chat now and then to e.g. add automatic git pushing including a nice svg chart of progress and to switch model strategies here and there i.e. to do a first pass with a cheap model and then switch to opus/sol if the small model can't solve it.

I could now one-shot a new game, yeah.

skeledrew•24m ago
The forgotten model is back on the map. I actually got OK mileage when I tried it for coding months ago. Maybe I'll try it again, with Opus guiding it, and see how it goes.
satvikpendem•22m ago
Apparently quite a bit smarter than Luna, I wonder what use cases it can cover. I actually honestly don't need a Haiku level AI to be that smart, and looks like you pay for it in the per token cost, I need speed mainly. I might even rather have a dumber but much faster model for things like web searching and parsing to retrieve results for the app or other LLM to do things with.
areoform•13m ago

    > but they still block penetration testing and other techniques more likely to be used by attackers.
    > 
    > Haiku 5.5’s biology safeguards are the same as for Sonnet 5, Sonnet 5.5, and Opus 5. They allow research biology questions but restrict access to requests that we judge as likely to cause harm. Organizations working on wider-ranging biology and cyber activities can apply to our Life Sciences Verification Program and Cyber Verification Program.
I would like to take a moment of your time to tell you about some of the "bioweapons" Anthropic has blocked that involved Haiku!

These are the examples from "Detecting and countering misuse of AI: September 2026" - https://news.ycombinator.com/item?id=49647300

    > Importantly, because our biological safety classifiers robustly block content involving high-risk biological research (in this case, the construction of enhanced pandemic potential pathogens), all of these exchanges occurred on models in our weakest class of models (specifically, the models were Claude Sonnet 4 and Haiku 4.5, the latter of which the user began using after Sonnet 4 was deprecated). 
    >
    > Upon a detailed examination of the exchanges, we estimate that the uplift provided by Claude was primarily clerical assistance in data analysis, study ideation and design. This is consistent with our understanding of the capabilities of Sonnet 4 and Haiku 4.5, which are not able to perform expert-level biology research tasks; we estimate that the uplift provided to the researcher was limited and substantially lower than it would have been from one of our more capable models.
Anthropic then says for the above, "we estimate that the uplift provided by Claude was primarily clerical assistance in data analysis, study ideation and design"

While doing my best to avoid comment, please note, they're talking about a domain expert in a state research institution using Claude to do paperwork.

What did they save us from? What bioweapons did these filters prevent? From the front matter report,

    > The above LLM platform is not the only route via which researchers engaged in viral gain-of-function research have used our platform. In May 2026, we discovered a researcher outside the US using Claude in their research on highly-pathogenic avian influenza (“bird flu”). The research focused on viruses’ adaptation to mammals, and the mechanism by which it causes severe disease beyond the respiratory tract.
OK. Sounds serious. "Gain of function research..." but who and why?

    > The researcher pursued this work in a credible institutional context, and interacted with Claude over the course of several weeks, exchanging thousands of messages. In these exchanges, the researcher leveraged Claude’s knowledge of the scientific literature to assist the researcher in study planning and design, data analysis, and the interpretation and prioritization of experiments. The researcher also used Claude for editorial assistance in writing up the research.
So this was a researcher inside of some country's national lab ("credible institutional context") doing research on dangerous viruses using Claude for "for editorial assistance in writing up the research."

What "uplift" are you providing to scientists working at specialized global BSL-4 labs that already have – and I quote their report - "physical access to such isolates." (as in samples of viruses)?

This nonsense has been expanded with even worse "safeguards."

This makes Claude unusable for any serious scientist or people curious about science, which is sad.

Topfi•10m ago
131tok/s P50 according to OpenRouter currently, though might move up or down over the coming days. If it sticks at that speed, roughly twice the throughput of Luna is impressive, though the 5x price increase beyond 100k is painful.

Was a big fan of Haiku 4.5, though understand why for most Sonnet was the far better option.

johnisom2001•10m ago
It fails the "How many r's in <word>?" test.

I ask:

> how many r's in diminished

It answers:

> Diminished has 1 r.

j45•49m ago
It could be to incentivize people to not be lazy users of tokens.
Eridrus•47m ago
It's actually existing flat per-token pricing that is weird.

Neither encode nor decode are linear in compute, so providers need to price for average expected length.

This is just getting closer to the true cost of generating tokens.

insanitybit•45m ago
I mostly use Haiku for really, really basic stuff, never for actual engaging work. I've used it for first-pass analysis to triage bugs, for example - all it does is related N bugs together to see if any potentially relate. Then I have Sonnet investigate further.
dannyw•45m ago
Haiku 5.5 is noticeably smarter than GPT-6 Luna, so I can see their pricing strategy here.

For a while Anthropic has lacked a cost effective “cheap” LLM for summarisation, compacting, RAG helpers, etc.

These ‘ephemeral’ workloads are often under 100k tokens, or can be structured to be under 100k.

In some coding benchmarks, Haiku 5.5 beats Sonnet 5! (Especially implementation; do a well defined Jira ticket; etc), it’s really impressive how much intelligence per dollar has grown in just a few short months.

system2•45m ago
Who in their right mind would use haiku while Mimo or GLM cost 10% of what they are charging with much smarter models?
user43928•41m ago
Presumably everyone who doesn't bother integrating a third party API key into their harness, which would probably be most of the Claude Code users.
wyrdcurt•16m ago
Some people/organizations are ideologically opposed to using Chinese models. Not me, I use GLM-5.3-Flash almost everything (the subscription-subsidized pricing on a legacy Z.ai plan makes it the best value model by a wide margin), along with some MiMo and DeepSeek. Still, I use Luna for certain tasks where speed is more valuable than performance; I can see this new Haiku displacing Luna for those. If you mean Haiku 4.5 though I agree, that model was a waste of time and money.
mrngld•4m ago
That's not what any benchmarks that look at cost per task or similar says in terms of cost. The Chinese models, generally speaking, might be cheaper per token but need a lot more tokens to get there.
esafak•45m ago
It's their creative way of 'matching' Luna's prices.
AustinDev•44m ago
encode and decode tok/s which is ($/s) when it comes to pricing drops heavily above 100k tokens.

There are plenty of workflows like translations where you'd easily be under the cap.

enraged_camel•44m ago
>> 100k tokens is an absurdly low cutoff and it is only applicable to Haiku and not Sonnet or Opus. It's a low enough cutoff that it will be quickly exceeded if you are doing anything with Agents

Your vibes don't appear to be supported by facts. From the announcement:

>> Claude Haiku 5.5 is priced 90% lower than Claude Haiku 4.5 for requests up to 100,000 tokens, and 50% lower for requests over 100,000 tokens. On Haiku 4.5, 90% of requests fell into the former category.

Philpax•41m ago
People weren't using Haiku 4.5 for agents before. 5.5 is good enough that it might be.
StilesCrisis•5m ago
Haiku 4.5 users were using it for Kleenex requests because that was the best it could do.
Tiberium•43m ago
There's also a tokenizer efficiency difference: modern Claude's 100K tokens are about ~60-65K modern GPT tokens, so in reality the Luna cutoff is much further away than the Haiku one.

You can test with Anthropic's count_tokens endpoint or with https://crates.io/crates/tokwc

AtNightWeCode•24m ago
> ...this tokenizer, the same input text produces approximately 30% more tokens on Claude Haiku 5.5 than on Claude Haiku 4.5.

So, it is might be even worse.

Tiberium•20m ago
No, it's just Haiku 4.5 is so old that it predates the new Claude tokenizer change in Claude 4.7+
HarHarVeryFunny•38m ago
Notable that one suggested use case for Haiku is "classification requests", i.e. Jev competitor, and the pricing matches GPT-6 Luna which is behind OpenAI's "Decisions API" Jev competitor.

For this application 100K token input is plenty.

Of course Anthropic and OpenAI, both at $0.10/M, are still 2.5x the cost of Jev's $0.04/M.

martianvoid•34m ago
I think the 2.5 times cost but actually pays off in terms of intelligence compared to jev and the general capability of using it beyond classification
HarHarVeryFunny•20m ago
The classification performance remains to be seen, but presumably we'll soon start to see classification benchmarks.

For other tasks like summaries (another suggested usage) it's good to see Luna and Haiku now competing against each other on cost.

I'd love to know how the business automation market breaks down by volume of call type though - hard to imagine that decision making (e.g. branching, triage) isn't a very large part of it, greater than these other suggested Haiku use cases.

port3000•28m ago
They are targeting businesses/API use for fast decision making and agent integration. Plus they now need to be competitive with Jev-type models in that space.
swalsh•41m ago
I've been using GPT-6 Luna in some capacity for nearly all my agent workflows. It's just a really good model, and the pricing is cheap. If Haiku 5.5 is better, and the same price (under 100k context... which is a big caveat) i'd probably swap it.
dannyw•36m ago
It’s absolutely better than Luna. It feels closer to a “sonnet 5.2” if that makes sense.

Of course it’s not as big, and hence falls-off quicker. I’d consider the 100k a “promotional price” to match Luna’s token pricing while delivering noticeably more intelligence.

Plutoberth•29m ago
I'm building a game that incorporates LLMs as a game mechanic.

I've been using Luna, but I'll probably switch to Haiku.

notatoad•13m ago
not haiku, but luna - last week i used it for things like "read this historical dump of 15k support tickets and break them into categories that make sense, then propose help docs that i could write to handle the most frequent queries in each category"

used <10% of my 5hr limit on a $100 codex plan.

•
17m ago
That's an old tactic for an old world. You only need, what, half an hour with your agent of choice to write you out of that?
losvedir•12m ago
Nah, it's pretty trivial to switch providers (especially with Claude's help, ha).

This is more to encourage people to try out adding AI into their product, which is a totally different flow and experience from using AI to build the product.

tech234a•31m ago
OpenAI will probably add this to their plans within a week
alasano•18m ago
With OpenAI you can just use Oauth and get a token to use your subscription.

Anthropic isn't even close to being this useful.

Iolaum•9m ago
Biggest reason for an OAI subscription instead of Ant imo.

Biggest loss is that Ant models look like they are genuinely better.

thepasch•14m ago
This is them sneaking in taking the Claude Agent SDK off of subscription plans through the back door along with a model release. They previously wanted to do this in June, but backpedaled after huge backlash:

https://support.claude.com/en/articles/15036540-use-the-clau...

•
15m ago
Totally fair, but I'd encourage you not to look at the design so much as the task. This was a design that's part of a benchmark test suite specifically for image->html conversion. The dense visual noise / complexity / flowing svg shapes are things that most LLMs have trouble with.

It's meant to be a good test, not a good design.