frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Audacity 4.0

https://github.com/audacity/audacity/releases/tag/Audacity-4.0.0
199•ClydeN•1h ago•38 comments

Pre-Release of Polars 2.0

https://pola.rs/posts/announcing-polars-2/
244•komape•5h ago•71 comments

The Browser's Main Thread Is Expensive

https://kciter.so/posts/the-expensive-main-thread/en/
180•kciter•1d ago•51 comments

Invisible Companies

https://colossus.com/article/invisible-companies/
23•ltononro•2d ago•5 comments

Muse Spark 1.3

https://developer.meta.com/ai/models/muse-spark/
620•bvaldivielso•17h ago•406 comments

Gemini 3.8 Flash and 3.8 Flash Cyber

https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-c...
1065•bratao•21h ago•599 comments

What I Learned from My Mom (1941-2026)

https://experimentalliving.substack.com/p/what-i-learned-from-my-mom-1941-2026
66•NaOH•4d ago•4 comments

9 Mothers (YC P26) Is Hiring in Austin, TX

https://9mothers.com/careers
1•ukd1•50m ago

Fish Bad, Sugar Good and Other Medieval Ideas About Food

https://lithub.com/fish-bad-sugar-good-and-other-medieval-ideas-about-food/
30•mooreds•2d ago•19 comments

Three schoolgirls in Kinsale pulled up a pea plant covered in warts (2016)

https://scienceblog.com/b-three-schoolgirls-in-kinsale-pulled-up-a-pea-plant-covered-in-warts-and...
99•DamonHD•5h ago•37 comments

Three sites made 215,128 “best software” pages for AI. Perplexity cites them

https://trellner.com/reports/manufactured-sources-behind-ai-recommendations/
455•jakobgreenfeld•22h ago•220 comments

Google avoids a breakup of its ad tech business

https://www.nytimes.com/2026/09/02/technology/google-ad-tech-remedies.html
418•donohoe•22h ago•296 comments

The Computer Museum of America reclamation project

https://computer-museum.org/wp/
62•rbanffy•2d ago•27 comments

Holden's Lightning Flight

https://en.wikipedia.org/wiki/Holden%27s_Lightning_flight
178•ColinWright•3d ago•39 comments

Intrusive Linked Lists

https://www.data-structures-in-practice.com/intrusive-linked-lists/
6•tripdout•3d ago•0 comments

Can I opt out of my input or output data being used for training?

https://help.mistral.ai/en/articles/455207-can-i-opt-out-of-my-input-or-output-data-being-used-fo...
467•teekert•1d ago•215 comments

Google Antigravity TOS: 3rd party usage can get Google account suspended

https://twitter.com/GergelyOrosz/status/2095453567955968398
24•tosh•1h ago•4 comments

A dark horse enters China's AI race: StartLux

https://chinaonchina.com/article/chen-dawei-returns-enters-the-large-model-sector
20•try-working•1h ago•3 comments

Fable 5.1 World Modeling

https://github.com/PhiloLabs/fable51-worlds
285•surreal_•17h ago•82 comments

Reverse Engineering Unknown File Formats with ImHex

https://werwolv.net/posts/file_format_reverse_engineering/
222•carlos-menezes•3d ago•42 comments

Higher Multipoles of the Cow

https://arxiv.org/abs/2504.00506
94•MrOrelliOReilly•2d ago•25 comments

Launch HN: RonanRX (YC S26) – Personalized Peptides and GLP-1s

72•lloydarmbrust•14h ago•74 comments

Biggest dark matter detector spots a single weird particle

https://www.science.org/content/article/world-s-biggest-dark-matter-detector-spots-single-weird-p...
312•randycupertino•23h ago•112 comments

Aging brains blend memories together instead of just forgetting them

https://studyfinds.com/aging-brains-blend-memories-together-instead-of-forgetting-them-study-finds/
303•mdp2021•23h ago•123 comments

Wendell Berry has died

https://www.nytimes.com/2026/08/31/us/wendell-berry-dead.html
206•Curiositry•2d ago•101 comments

Qantas Airbus A380 engine failure in 2010 (2023)

https://admiralcloudberg.medium.com/a-matter-of-millimeters-the-story-of-qantas-flight-32-bdaa62d...
159•gumby•18h ago•92 comments

Engineering of the fastest WebAssembly interpreters

https://wasmi-labs.github.io/blog/posts/wasmi-v2.0/
116•herobird•2d ago•14 comments

Exit the Cave

https://turtlespace.blog/p/exit-the-cave
284•akkartik•22h ago•97 comments

A Selection of Los Alamos Rolodex Business Cards

https://clui.org/collections/los-alamos-business-cards/selection-cards
190•1970-01-01•2d ago•50 comments

Poisson Disk Sampling

https://stripeacross.com/posts/poisson-disk-sampling/
180•vismit2000•23h ago•21 comments
Open in hackernews

4.5B Posts Scraped from TikTok

https://tiktok-api.seeksocial.io/
40•TheOnlyWayUp•1h ago

Comments

smallerize•51m ago
There's no way this dataset is going to survive on HF, right? It will be hit with so many DMCA takedowns.
deviation•50m ago
My thoughts also
simonw•9m ago
It's metadata only, not video content. I expect it will likely survive - that's a pretty common pattern for machine learning datasets.

Stable Diffusion was enabled by LAION, for example. That was metadata about images and URLs to those images, but not the actual image files.

ckugblenu•50m ago
This is being posted all over the place. on multiple subreddits and stuff. Why?
vachina•26m ago
Ads for an exploit
Retr0id•23m ago
Replicating client behaviour is hardly an exploit.
Retr0id•49m ago
Can any brave soul wade through the LLM prose to provide a human-readable summary?
405error•41m ago
I cannot verify whether it is technically correct, but it's about how to defeat Tiktok's bot filters to scrape it.
arcfour•35m ago
Man, I could use $699. If this gets even a single sale, then maybe I need to try having less decency...

(I guess that is sort of a roundabout summary...)

echoangle•32m ago
Would you say that this is immoral? Since it’s data that’s public anyways, I don’t see why putting it in a table and selling it is a bad thing.
arcfour•28m ago
No, I don't care at all, but I recognize that my morals might be lower than others here. To me it's just public data...whatever.

My criticism was basically - this is trying to sell an AI slop project for $699 a pop - I could get this out of a few Claude Code sessions if I had the storage and network bandwidth to run such a scraper. The value proposition is questionable when the writing shows that the entire project was AI generated, and clearly Claude understands the way the TikTok Android app internal API works quite well...

arvid-lind•22m ago
igor_nast•48m ago
Is this even useful in any way?
magicmicah85•47m ago
>Is it legal? It is against TikTok's terms of service. It is sold for research and educational use.

Oh, ok. Otherwise, very detailed deconstruction to scrape their API. Lots of layers of registration and creating a request that looks like it is valid client.

pr337h4m•46m ago
There aren't any actual videos in the dataset though.
405error•36m ago
The first immediate smell is that if you have 4.5B rows and 289GB in data, you have ~60 bytes per row.
smallerize•19m ago
The data is listed near the bottom, under "The 24 Endpoints" https://tiktok-api.seeksocial.io/#api
gdevenyi•45m ago
Like gazing into hell
VulgarExigency•36m ago
In 2026, everything we say, see or do on the internet is just another data point for some unscrupulous bastard to seek profit from.
igor_nast•19m ago
That's not true. I build an agentic IDE - indie dev tool. And it's for free and stays like it. I do it because I enjoy coding and the technology :)

https://shikigami.dev - more details here

raver1975•40m ago
That data does not belong to the public. Why do you think it is OK to steal from TikTok?
theplumber•36m ago
Is this sarcasm on AI companies?
moinism•39m ago
> Everything described here is a private Go repository. One-time payment, permanent access, complete source. > $699 one time · lifetime access

Not open-source apparently.

And I cant find the reddit post but I think I read that videos/assets are not actually pre-downloaded, they have to be requested through Tiktok API using the provided code. So if Tiktok patches, the code will need updates too.

nomilk•32m ago
I think the post is essentially a decent technical-explainer (value adding and interesting) in exchange for effectively a little product placement (selling either just the code or code already running on a server at additional cost)

But I think this is the 289GB data (free): https://huggingface.co/datasets/kuben-developer/tiktok-video...

VulgarExigency•37m ago
Absolutely no consideration for the people who, when posting things to Tiktok, would prefer for them not to get scraped.
nomilk•29m ago
> Three things to notice, because each one bites later:

Very LLMish language!

405error•23m ago
It's that mix of dense, impressive sounding jargon, but even scanning across it raises glaring problems. Like, if you have 4.5 billion videos on HF, and it's 289GB, it's about 60 bytes per video. Checking the column fields as well, there doesn't seem to be any video files*.
nomilk•17m ago
There's an `is_video` column, perhaps containing a lot of 0's
405error•11m ago
More directly, there simply aren't any video files uploaded. It's just parquet files, which contain no video columns (I'm not even sure if it supports it).
vachina•11m ago
The AI keeps mentioning how a HTTP 200 can silently pollute your dataset. Why not just check contents of body? Usually APIs follow strict JSON contract for successful queries, alert or throw an error when that changes.
405error•5m ago
It's probably AI coded and hallucinated many things. That drumming up the importance of a minor thing is a real tell. Another hallucination - it hasn't found any video APIs (despite statements that it has and uploaded it). It has video metadata.
The AI giants got where they are by pushing similar limits and setting aside moral (and legal) issues to be settled later in court, so the idea seems to fit the general zeitgeist we're living in. Not really criticism, just an observation.
breezybottom•28m ago
What is "public" data? It admits to violating the ToS.
arcfour•25m ago
Something posted with the understanding that it could potentially be viewed by anyone on the Internet with minimal/no restriction.

ToS are just what you follow if you don't want to get banned off of the site. If you don't care about that, then you can go hog wild, though you're being a bit of a jerk/not playing nice obviously.

breezybottom•22m ago
Any company could potentially be hacked, so by that standard all data is public. And no, the ToS is legally binding on the company as well.
arcfour•19m ago
That's a blatant strawman - that's not a reasonable position at all. A reasonable person does not expect their private medical records to be accessible to you or me just because they are stored in an EHR system.

They might be surprised that you or I looked at their TikTok video when we aren't the intended audience, but they still posted it publicly, with the understanding that it would be made freely available to others.

breezybottom•13m ago
It's definitely not a strawman. Perhaps you meant that it's a false equivalency, but I don't think that's true either.

With how common data hacks are, why wouldn't a reasonable person expect their medical records to leak? I received at least two such breach notices just last year.

simonw•11m ago
"I uploaded 4.5 billion of those videos to Hugging Face" = I scraped the metadata (title, view count, etc) for 4.5 billion TikTok videos using the same API as their Android app. I posted the data to HF as Parquet. The video content itself is not included. I'll sell you my Go scraping code.