frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Cloudflare acquires Deno

https://deno.com/blog/cloudflare
751•ilreb•5h ago•399 comments

Our $445M Series D

https://oxide.computer/blog/our-445m-series-d
408•ahlCVA•5h ago•165 comments

Sorry, I'm in a meeting

https://iminafleeting.com/
508•splintersio•8h ago•171 comments

Show HN: Let your AI agents paint big arrows, boxes and text on your screen

https://github.com/franzenzenhofer/big-arrow-on-the-screen
310•franze•7h ago•131 comments

Nobel Peace Prize for 2026 to Navanethem Pillay

https://www.nobelprize.org/prizes/peace/2026/press-release/
345•Anon84•8h ago•172 comments

Germany transforms former coal mines into Europe's largest lake landscape

https://www.euronews.com/2026/04/14/almost-like-lake-como-germany-transforms-former-coal-mines-in...
122•ohjeez•3h ago•62 comments

Training Text-to-Image Models Without a VAE

https://www.linum.ai/field-notes/pyramid-jit
16•schopra909•2d ago•7 comments

Why isn't the industry freaking out about DeepSeek 4.1 Flash?

https://www.dgt.is/blog/2026-10-07-deepseek-freek-out/
1002•jonotime•1d ago•906 comments

Whistle: Speech to Text in 16.9 MB

https://cactuscompute.com/blog/whistle
886•gmays•1d ago•174 comments

Man discovers his parents' coffee machine used 1TB of data in 10 days

https://www.dexerto.com/entertainment/man-discovers-his-parents-coffee-machine-used-1tb-of-data-i...
891•ck2•2d ago•536 comments

Triple-A Minesweeper

https://minesweeper.mikelacher.com/
26•robin_reala•2h ago•4 comments

I hired an illustrator to draw my house. Now it's my Home Assistant dashboard

https://antonfrolov.substack.com/p/i-hired-an-illustrator-to-draw-my
844•soheilpro•2d ago•184 comments

MXC - a sandboxed code execution system

https://github.com/microsoft/mxc
142•nreece•12h ago•70 comments

Yes, and

https://htmx.org/essays/yes-and/
666•Michelangelo11•1d ago•267 comments

'Wallace and Gromit,' 90% Alone

https://animationobsessive.substack.com/p/wallace-and-gromit-90-alone
20•vinhnx•4h ago•0 comments

Once: Cache CLI commands

https://github.com/alex0ptr/once
74•baquero•8h ago•33 comments

Scam American companies are using to manipulate ingredient lists

https://twitter.com/WallStreetApes/status/2108594998656807078
20•bilsbie•42m ago•10 comments

Keyboard differences between Windows and Macs

https://unsung.aresluna.org/deeper-dive-keyboard-differences-between-windows-and-macs/
276•sohkamyung•15h ago•227 comments

Theranos.world

https://www.theranos.world/
528•kbyatnal•1d ago•189 comments

OpenAI fires three safety researchers for "mishandling research information"

https://techcrunch.com/2026/10/08/fired-openai-safety-researchers-dispute-misconduct-claims-warn-...
230•trakkstar•8h ago•141 comments

Programming Isn't Special

https://blog.glyph.im/2026/10/programming-isnt-special.html
103•ingve•10h ago•102 comments

100+ reactions to 100+ solutions

https://proofsandprompts.com/2026/10/08/100-reactions-to-100-solutions/
30•jacobedawson•5h ago•16 comments

Python 3.15

https://www.python.org/downloads/release/python-3150/
236•ngoldbaum•3h ago•51 comments

Ask HN: What do you run on a $5 VPS that's worth keeping online 24/7?

355•mariocesar•2d ago•570 comments

The value of not getting to the point (2015)

https://ken.arneson.name/2015/11/the-value-of-not-getting-to-the-point/
215•NaOH•23h ago•74 comments

The Hetzner Cloud network stack – history and technical overview

https://www.hetzner.com/blog/the-hetzner-cloud-network-stack-history-and-technical-overview/
84•cnkk•5h ago•23 comments

OpenAI, the Partition Principle, and Mathematics

https://karagila.org/2026/openai-pp/
182•md224•18h ago•274 comments

Iranian campaign planted fake articles in real U.S. publications using ChatGPT

https://www.washingtonpost.com/technology/2026/10/09/chatgpt-users-iran-planted-ai-generated-arti...
157•cainxinth•5h ago•137 comments

ETH-68: Ethernet Audio Interface for Linux

https://naturalsystems.io/eth68
206•chabad360•2d ago•131 comments

DuckDB Ducklake

https://github.com/duckdb/ducklake
226•saikatsg•2d ago•32 comments
Open in hackernews

Training Text-to-Image Models Without a VAE

https://www.linum.ai/field-notes/pyramid-jit
15•schopra909•2d ago

Comments

schopra909•2d ago
Hi HN, author here!

For context, we're a 2-person lab training generative video models. Goal is a new set of controllable, animation tools (you can read more about that here https://www.linum.ai/about if you're curious).

The biggest bottleneck for our last text-to-video model in terms of training and inference cost is attention. Video models are incredibly token dense (e.g. 110K tokens for a several second clip). If we can condense that context window more aggressively, we can train bigger models for a lot less $$ and offer them to prosumers at reasonable price points (unlike the big models today like Seedance, which cost an arm and a leg to run).

Traditionally, image and video models have two disjoint components: VAE (Variational Autoencoder) and Diffusion Transformer (DiT). They're trained separately, and empirically VAEs seems to struggle to get past 16x16 token reduction.

Here, we're switching to pixel-space, throwing away the VAE, and achieving 32x32 token reduction (4x smaller context windows) while learning a better overall model in a fraction of the training samples.

The central thesis is "simpler is better". If we can put the compression problem into the more powerful Diffusion Transformer (DiT) would should be able to learn a "latent space" optimized for generation and get better compression without hurting generation quality.

I'll be checking this post off and on the next couple of hours, so feel free to drop questions below. And I'll try to answer them to the best of my ability.

P.S. The model checkpoints from this blog are Apache 2.0, so feel to try playing with it yourself on a GPU!

E-Reverance•2d ago
I know it goes a tiny a bit against the spirit of what y'all are doing, but applying a few layers of pixel-wise local attention (so 1x1 "patch", with 3x3 or 5x5 attention window, basically treating it as a dynamic conv) has worked way better than both linear and conv unpatching in my recent experiments.

Diagram for reference https://x.com/1rreverant/status/2107546198093730287 (In my most recent recent experiment I actually removed the MLP and just used a linear project on the pixel's hidden states)

schopra909•2d ago
That sounds like an interesting idea!

Can you confirm I'm understanding correctly?

1) Linear unpatchify as usual to go from hidden states to pixel space

2) Attention within a local window (e.g. 3x3, 5x5) to "blend" pixel space data and come up with a better image (as an alternative to MLP or Convolution)

And follow up questions:

1) How do you handle boundaries between your "attention windows"? Do you move the window just like a convolution does or are the "attention windows" all mutually exclusive from one another?

2) How much faster/slower is this operation vs. a linear layer + MLP?

E-Reverance•2d ago
1) yes, but with more than channels than 3 (and no correspondence to color, its "pixel space" in the sense of position, not value)

1D grid example for clarity (obviously meant to be done in 2D though)

So embed -> [N dim, N dim, N dim ... , N dim] instead of embed -> [RGB, RGB, RGB ... , RGB]

2) / [followup 1)] Stride of 1, so we place a window at each pixel (so lots of overlapping)

Also I wouldn't phrase it as "come up with a better image", the point is to give less spatial decoding pressure to the patch tokens so that they can almost completely focus on feature learning instead. There is no reason to have the model learn spatial decoding when the structure prior of images is comically strong (especially compared to text), its a waste of training time and parameters

[followup 2)] I haven't measured but it was passable is all I can say (my experiments setup are horrendous right now lol)

(Also don't mind the phrasing, I just wanted to be 100% clear)

bitpush•33m ago
Is there a way to use this in ComfyUI today?
schopra909•2d ago
I appreciate the clarifications!

If we have extra compute lying around in the coming weeks, we’ll try this out and report back. It’s a good idea :)

E-Reverance•2d ago
I loved to hear! Just ping me on twitter (same account I sent the link with)