frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Show HN: Offlineisbetter: Efficient Models for Text on CPU

https://github.com/offlineisbetter/offlinedemo
1•chronicallyoffl•40m ago

Comments

chronicallyoffl•40m ago
hi folks,

i'm an engineer and amateur hacker who's been working on a little side project, and i want to show it off and also solicit your feedback.

i'm building efficient machine learning models for natural language processing on consumer cpus: sentiment analysis, task routing, document retrieval, etc. i know there are a lot of models for these tasks already, but too many of them are:

(1) difficult to configure, or (2) parameter-inefficient, or (3) too high-latency for most applications.

instead of building academically "novel" architectures, my project is making parameter-efficient and low-latency models easy to use through distillation or finetuning, computation graph optimization, and quantization. the eventual vision is a single terminal command to install the model. (i've not quite achieved this, it's like three or four commands right now).

i'm posting here because i've completed a beta version of my first model and i'd be very grateful for community feedback. the first model is called `offline-sentiment-small` and it's for binary sentiment analysis (positive/negative). you can install the runtime using `pip install offlinedemo` and download the checkpoint from the releases page of

https://github.com/offlineisbetter/offlinedemo

the readme of the above repository shows benchmarks against three other models for sentiment analysis: distilbert, roberta, and modernbert. they all do about the same on sst2 (~0.92-0.95 f1 score) because it's a relatively easy dataset. you'll notice that, though my model is 230m parameters, you need to go down to distilbert (67m) to get the same kind of latency. i want to stress-test these models on datasets with longer-context or more difficult sentiment tasks, to see how performance degrades with both my model and these models. if you know about any good datasets for this, please do let me know!

for those curious, my model is a liquid foundation model (lfm2.5) encoder (on hf, LiquidAI/LFM2.5-Encoder-230M) finetuned by low-rank adaptation on the sst2 training set, optimized with onnx, and quantized to int8. i chose lfm over bert as the base model because it uses grouped-query attention and convolution layers instead of dense attention. these changes over the "vanilla" transformer architecture make its latency subquadratic in the input context (and i've verified this empirically). lfms are often considered more parameter-efficient than transformers: for example, liquid ai's 2.6b autoregressive model competes with llama 7b.

and disclaimer, i have no affiliation to liquid, i just think their work is cool!

i want to emphasize that i'm not claiming technical novelty! rather, i'm building a way to make parameter-efficient, distilled/finetuned, and quantized models more accessible to more people. right now, if you wanted to run a distilled and quantized text model, it takes a non-negligible amount of compute and time and effort to set it up. not everybody wants to do that, and so out of convenience they'll just opt for a big llm in the cloud.

i worry that so many people these days pay for openai/anthropic tokens just to do a simple task, such as sort their emails into categories. not only is this wasteful, high-latency, bad for the environment, etc. but you're giving somebody else your data. so i've started the offlineisbetter project to help raise awareness to wasteful model use and provide the community with an alternative.

please let me know if you have issues testing the model on your machine. i'd be very grateful to hear any feedback you might have!

Show HN: Karma Compass – Agentic board management for nonprofits

https://compass.karmahq.org
1•mmurthy•34s ago•0 comments

Every Pho Number

https://everyphonumber.com/
1•c249709•57s ago•0 comments

Show HN: Cascade Screening, a local sanctions screener with published test sets

https://github.com/ArslaneSempai-ui/cascade-screening
1•arslanechr•1m ago•0 comments

THB #882: Digger (Spoilers)

https://davidpoland.substack.com/p/thb-882-digger-spoilers
1•layer8•1m ago•0 comments

Inspect: An open-source framework for large language model evaluations

https://inspect.aisi.org.uk/
1•erikcw•2m ago•0 comments

Show HN: We built a memory layer for web agents and made them 3x faster

https://blog.reduck.ai/browser-action-memory-layer-for-faster-agents/
3•Labo333•2m ago•0 comments

Show HN: Ikarem – Zero-dependency Python ASGI framework with built-in MCP tools

https://github.com/nishantXnova/IKAREM
1•NishantPaudel•3m ago•0 comments

Movements in 3D

https://507movements3d.com/
2•dllu•3m ago•0 comments

I bought a second iPhone to use my first one less

https://edparry.com/the-play-phone
1•edparry•4m ago•0 comments

LZ in OpenZL

https://openzl.org/blog/2026-09-29-lz-in-openzl/
1•terrelln•5m ago•0 comments

Pilot of Israel-bound FlyDubai flight tried to crash plane

https://www.ft.com/content/7ed6ed48-79e4-46c2-968f-f2a69d263d35
1•alephnerd•5m ago•0 comments

Pilot arrested after stabbing anther pilot and trying to crash Israelbound plane

https://www.theguardian.com/world/live/2026/sep/30/dubai-israel-saudi-arabia-flydubai-flight-dive...
3•Markoff•8m ago•0 comments

You'll never reach 85% on BEAM 10M using .md files

https://past.dev/blog/md-files-vs-memory-api-beam-10m
5•mehdidjabri•11m ago•1 comments

Minitel

https://en.wikipedia.org/wiki/Minitel
1•GuinansEyebrows•11m ago•0 comments

There is no Percolation at the Critical Probability in all Dimensions

https://gilkalai.wordpress.com/2026/09/03/amazing-there-is-no-percolation-at-the-critical-probabi...
1•abetusk•12m ago•0 comments

Diddy's pampered life in prison: Massages, nude photos, Hennessy

https://www.nbcnews.com/news/us-news/diddys-pampered-life-prison-massages-nude-photos-hennessy-rc...
1•speckx•12m ago•0 comments

DoorDash opens US waitlists for AI ordering agent and bulk-order API

https://about.doordash.com/en-us/news/doordash-ordering-connector
1•utiiiD•12m ago•0 comments

SynthID Bio: Watermarking methods for synthetic biology – Google DeepMind

https://deepmind.google/blog/introducing-synthid-bio/
1•xnx•13m ago•0 comments

Show HN: An old school forum on an open protocol – your identity is a keypair

https://nostr-proto.org/
1•dtonon•17m ago•0 comments

CoW – a stacking window manager for Wayland

https://cow-wm.codeberg.page/cow/
2•birdculture•18m ago•0 comments

Ask HN: Point of view about Agents Identity?

1•mathieu_aithos•18m ago•0 comments

Pledge signed by President Trump and top AI leaders misspells the United States

https://techcrunch.com/2026/09/30/pledge-signed-by-president-trump-and-top-ai-leaders-misspells-t...
7•CharlesW•18m ago•2 comments

Show HN: Bongard-mini – An open-weight Jev-like model built on T5

https://huggingface.co/AgentBull/bongard-mini
1•borisding1994•18m ago•0 comments

Show HN: Mingbird – an agent harness that makes a 2B model finish real tasks

https://github.com/Mingbird/Mingbird-agent
3•Gustor•18m ago•0 comments

Cities Forced to Funnel License Plate Data to Federal Surveillance Program

https://www.404media.co/how-cities-are-forced-to-funnel-license-plate-data-to-a-massive-federal-s...
2•CharlesW•19m ago•0 comments

Thoughts on AI

https://claude.ai/anthropic-interviewer/your-thoughts-on-ai
1•tzury•20m ago•0 comments

Forge Structures Requests for Prompt Cache Reuse

https://www.mohitranka.com/blog/how-forge-structures-requests-for-prompt-cache-reuse/
1•speckx•20m ago•0 comments

What control engineering taught me about building AI systems

https://gerbenrijpkema.substack.com/p/what-control-engineering-taught-me
2•gmrijpkema•21m ago•1 comments

Show HN: BreachProbe, find out if your app leaks its database

https://breachprobe.thecompound.tech
1•kyisaiah47•22m ago•0 comments

Giving everything away for free [video]

https://www.youtube.com/watch?v=L6ykiChEwpU
1•gmays•22m ago•0 comments