frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Open-Source Refact.ai Agent is #1 on SWE-bench Lite With a 59.7% Score

https://refact.ai/blog/2025/sota-on-swe-bench-lite-open-source-refact-ai/
3•kate_at_refact•1y ago

Comments

kate_at_refact•1y ago
Open-source Refact.ai achieves #1 on SWE-bench Lite with a 59.7% score. Our approach: fully autonomous Agent, no manual intervention needed.

How we did this:

• Prompt strategy: https://github.com/smallcloudai/refact/blob/swe-boosted-prom... • Claude 3.7 Sonnet as orchestrator • deep_analysis() tool (powered by o4-mini) for reasoning • Tool suite for repository exploration, code modification, and testing. Used dynamically based on task needs • One correct solution through iteration!

Autonomy = our core strength.

Refact.ai Agent completes the entire dev workflow independently: plans, executes, tests, self-corrects, and delivers a production-ready result. For each task, it made one multi-step run to generate a single correct solution, creating custom strategies rather than following rigid scripts.

You can read tech details on our SWE-bench approach: https://refact.ai/blog/2025/sota-on-swe-bench-lite-open-sour...

Your questions are welcome! Also, welcome to try Refact.ai Agent in VS Code and Jet Brains: https://linktr.ee/refactai

Why Do We Exist? With Hakeem Oluseyi [video]

https://www.youtube.com/watch?v=r5KmXSLFK7c
1•binyu•45s ago•0 comments

OpenAI scraps release of Astra 6.1 model over safety issues

https://www.washingtonpost.com/technology/2026/09/28/chatgpt-maker-openai-scraps-release-astra-61...
1•lisper•2m ago•0 comments

Holo4: Powering generalist computer-use agents

https://hcompany.ai/newsroom/holo4
1•acossta•4m ago•0 comments

Show HN: Corral kill every command your agent starts

https://github.com/Cardinal44/corral
1•CG144•5m ago•0 comments

OpenAI Says It Will Not Release Newest A.I. Model Over Safety Concerns

https://www.nytimes.com/2026/09/28/technology/openai-astra-safety.html
5•jbegley•6m ago•0 comments

Show HN:Jylus – give AI systems evidence from changing data.

https://jylus.ai/try
1•JoshJH•7m ago•0 comments

Knockoff: A browser extension that filters pseudo-brand junk out of Amazon

https://github.com/Shpigford/knockoff
1•hentrep•8m ago•0 comments

Show HN: Jauvex 1.2, two-way voice chat harness for Claude+Codex+Grok+Jev

https://github.com/reindent/jauvex
1•daraosn•9m ago•1 comments

AI Companies Are Not (Necessarily) Liable for Unintended AI Cyberattacks

https://sarahconstantin.substack.com/p/ai-companies-are-not-necessarily
1•theptip•11m ago•0 comments

What is the best LLM for turbo charging code?

1•ffgfghgffggxxxx•12m ago•1 comments

OpenAI Halts Model Release Amid Safety Escalation

https://www.techbuzz.ai/articles/openai-halts-model-release-amid-safety-escalation
1•hackernj•15m ago•0 comments

Its not just the f*cking sandbox

https://twitter.com/joedaroo/status/2104335929293127851
3•jumploops•22m ago•0 comments

AI and the Revenge of the Non-Techies

https://maroun-baydoun.com/blog/ai-revenge-non-techies/
1•maroun-baydoun•22m ago•0 comments

AI's next training data: your dodgy gaming skills

https://www.wired.com/story/the-next-evolution-of-ai-is-learning-from-your-dodgy-gaming-skills/
1•promptspheree•25m ago•0 comments

Humanos – Help Building the Human Operating System

https://tryhumanos.com
3•zevinclark•28m ago•0 comments

Making Things to Make Them

https://aimode.substack.com/p/making-things-to-make-them
1•warthog•29m ago•0 comments

GPT-6 Astra Is the Best Vision Model We Have Tested

https://blog.roboflow.com/gpt-6-astra-vision/
2•gmays•30m ago•0 comments

Why embeddings don't solve RAG [video]

https://www.youtube.com/watch?v=kBFsKVBDU8o
2•softwaredoug•32m ago•0 comments

LLM makes decisions to raise its training scores, and ignore user directives

https://joinhandshake.com/research/ai/deepswe-reward-hacking/
2•guardiangod•33m ago•1 comments

AMD Acquires Fei-Fei Li's World Labs for $8.2B

https://www.bloomberg.com/news/videos/2026-09-28/amd-acquires-fei-fei-li-s-world-labs-for-8-2-bil...
2•sbulaev•33m ago•0 comments

1996 chat room simulator connected to Win95 and System 7 web desktops

https://lolchat.rip/
4•henrychannel•33m ago•3 comments

Prenatal exposure to the plasticizer DEHP increases autism and ADHD

https://linkinghub.elsevier.com/retrieve/pii/S2666634026002941
2•OutOfHere•34m ago•2 comments

Atrocious AI-Written Tests

https://gruhn.me/blog/2026-09-29/
2•ngruhn•35m ago•0 comments

I thought I was building a C replacement. I was wrong

https://c3-lang.org/blog/i_thought_i_was_building_a_c_replacement/
7•birdculture•35m ago•0 comments

Tech Utopians Want to Build a City for the Post-A.I. World

https://www.nytimes.com/2026/09/22/world/americas/praxis-uruguay-colonia-tech-utopia.html
1•georgex7•35m ago•0 comments

Show HN

2•humbu-ndou•42m ago•1 comments

China broadens travel curbs to encompass family of top AI talent

https://www.business-standard.com/world-news/china-broadens-travel-curbs-to-encompass-family-of-t...
2•jnord•53m ago•0 comments

Show HN: LightSpeed – Interactive visualizer of time dilation from 0 to C

https://lightspeed.webland.pl/
1•webland•53m ago•0 comments

Reverse Engineering How Meta's Muse Shops

https://caeliai.com/blog/reverse-engineering-how-muse-shops
1•kalanpeace•56m ago•0 comments

Anthropic's IPO prospectus shows AI vision, surging costs

https://www.reuters.com/business/finance/anthropics-ipo-prospectus-shows-sweeping-ai-vision-surgi...
43•6thbit•59m ago•37 comments