frontpage.
newsnewestaskshowjobs

Made with ♥ by @iamnishanth

Open Source @Github

Open in hackernews

Create Missing RSS Feeds with LLMs

https://taras.glek.net/posts/create-missing-rss-feeds-with-llms/
2•alastairr•9h ago

Comments

PaulHoule•8h ago
The general story about the LLM-scraper problem is that (1) "companies like OpenAI run badly implemented web crawlers to get training data" but there is (2) with LLMs scrapers could do content understanding (inference) that would make them more useful and I think the even more impactful (3) LLMs will empower people to write scrapers that would never have written them before.

I kinda laugh at (3) because it's been a running gag for me that management vastly overestimates the effort to write scrapers and crawlers because they've been burned with vastly underestimating the effort to develop what look like simple UI applications.

They usually think "this will be a hassle to maintain" but it usually isn't because: (a) the target web sites usually never change in a significant way because UI development is such a hassle and (b) the target web sites usually never change in a significant way because Google will punish them if they do [1]

It is like 10 minutes to write a scraper if you do it all the time and have an API like beautifulsoup on your fingertips, probably 20 minutes to vibe code it if you don't.

I am still using the same HTML scraper to process image galleries today that I used to process Flickr galleries back in the 00's, for a while the pattern was "fight with the OAuth to log into an API for 45 minutes" or "spend weeks figuring out how to parse MediaWiki markup" and then "get the old scraper working in less than 15 minutes". Frequently the scraper works perfectly out of the box, sometimes it works 80% out of the box, always it works 100% by adding a handful of rules.

I work on a product that has a React-based site and it seems the "state of the art" in scraping a URL [2] like

   https://example.com/item/8788481
is to download the HTML and then all the Javascript and CSS and other stuff with no cache (for every freaking page) and run the Javascript and have something scrape the content out of the DOM whereas they could just go to

   https://example.com/api/item/8788481
and get the data they want in a JSON format which could be processed like item["metadata"]["title"] or just stuffed into a JSONB column and queries any way you like. Login is not "fight with OAuth" but something like "POST username and password to https://example.com/api/login with a client that has a cookie jar" I don't really think "most people are stupid" that often but I think it all the time when web scraping is involved.

[1] they even have a patent for it! people who run online ad campaigns A/B test anything, but the last thing Google wants is for an SEO to be able to settle questions like "will my site rank higher if I put a certain phrase in a <b>?"

[2] ... as in, we see people doing it in our logs

160k Impacted by Valsoft Data Breach

https://www.securityweek.com/160000-impacted-by-valsoft-data-breach/
1•Bender•46s ago•0 comments

Malicious .NET files conceal RATs in bitmap images

https://www.scworld.com/news/malicious-net-files-conceal-rats-in-bitmap-images
1•Bender•1m ago•0 comments

From the Debris of Halley's Comet

https://nautil.us/from-the-debris-of-halleys-comet-1208801/
1•Bender•2m ago•0 comments

LLMs Stability: Lyapunov Conjecture

https://hal.science/hal-04850283v1/document
1•northlondoner•2m ago•0 comments

Could Apple Exist Without Its Ties to China? Probably Not.

https://www.nytimes.com/2025/05/01/technology/apple-china-tariffs.html
4•bookofjoe•5m ago•1 comments

Acquisition made last year by Apple could lead to big AI announcement at WWDC

https://www.phonearena.com/news/acquisition-by-apple-could-lead-to-ai-powered-calendar-app-for-iphone_id170243
1•mikece•12m ago•0 comments

Uncomplicated-alert-receiver – show Prometheus alerts on heads up displays

https://github.com/jamesread/uncomplicated-alert-receiver
1•jamesread-hn•16m ago•1 comments

'A very exciting day': Canadian design may revolutionize the way ALS is treated

https://www.ctvnews.ca/health/article/canadian-designed-ultrasound-helmet-may-revolutionize-the-way-als-is-treated/
2•amichail•18m ago•0 comments

Show HN: Circuit Tutor– Chat with and make basic EE schematics via LLM

https://circuit-tutor.xyz/
1•andrewrn•20m ago•0 comments

This what it looks like when an explosion creates gold in space

https://www.cnn.com/2019/08/27/world/kilanova-gold-2016-scn-trnd
1•teleforce•23m ago•0 comments

Automate Q&A WhatsApp customer support

https://chromewebstore.google.com/detail/whatsapp-personal-assista/ilpmegkcflnmihommkgepbhpaofejdcl
1•second_comet•27m ago•1 comments

C++ and Rust: Different tools for the Job [video]

https://www.youtube.com/watch?v=usXhALmZI7Q
2•cod1r•28m ago•0 comments

Brace for Disorder as the Great Power Shifts Begin

https://www.ft.com/content/e45091ae-31c7-46b2-95bb-b8197655cd33
2•alephnerd•29m ago•0 comments

OneBit Adventure

https://www.onebitadventure.com/
1•Lwrless•29m ago•0 comments

AI Models Embrace Humanlike Reasoning

https://spectrum.ieee.org/chain-of-thought-prompting
1•pseudolus•32m ago•0 comments

Hydrostatic Paradox

https://www.scubageek.com/articles/wwwparad
1•mhb•32m ago•0 comments

Charles Bukowski, William Burroughs, and the Computer (2009)

https://realitystudio.org/bibliographic-bunker/charles-bukowski-william-burroughs-and-the-computer/
7•zdw•36m ago•1 comments

Mantra to Burn $160M OM Tokens, 50% From its founder, Following 90% Price Crash

https://www.coindesk.com/markets/2025/04/22/mantra-to-burn-160m-om-tokens-50-from-daos-founder-following-90-price-crash
1•PaulHoule•40m ago•0 comments

Show HN: Amotions AI – emotionally intelligent real-time AI sales teammate

https://www.amotionsinc.com/
1•xupianpian1•41m ago•0 comments

A man who tried to murder a 600-year-old tree (2020)

https://www.cnn.com/2020/08/30/us/treaty-oak-gbs-great-big-story-trnd
1•1659447091•43m ago•0 comments

What is wallet encryption and why it matters

https://www.cryptocrafted.org/crypto-software-wallet-cryptocurrency-security/what-is-wallet-encryption-and-why-it-matters
1•baligako•44m ago•0 comments

Brandon's Semiconductor Simulator

https://brandonli.net/semisim/
16•dominikh•44m ago•1 comments

Lix 2.93 "Bici Bici"

https://lix.systems/blog/2025-05-06-lix-2.93-release/
1•todsacerdoti•46m ago•0 comments

Bruno Simon: Portfolio website as a 3D car course

https://bruno-simon.com/
2•gaws•55m ago•0 comments

Cosmos-482 descent craft re-entry prediction

https://blogs.esa.int/rocketscience/2025/05/07/reentry-prediction-soviet-era-venera-venus-lander-cosmos-482-descent-craft/
2•_kb•1h ago•0 comments

Show HN: SkipCut – watch YouTube ad free on Any device, No install, No login

https://skipcut.com/
2•Nina_F•1h ago•0 comments

WebGL Water (2010)

https://madebyevan.com/webgl-water/
42•gaws•1h ago•11 comments

OSU OSL's Path to Sustainability – Call for Smart Solutions and Enduring Support

https://osuosl.org/blog/osl-future-update/
1•pabs3•1h ago•0 comments

Stability solution brings carbyne form of carbon closer to practical application

https://phys.org/news/2025-05-stability-solution-unique-carbon-closer.html
1•westurner•1h ago•0 comments

Nearly 200 students under quarantine outbreak most infectious disease

https://www.dailymail.co.uk/health/article-14696291/students-quarantine-measles-outbreak-north-carolina-college-campus.html
6•Bender•1h ago•1 comments