frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Open-Source Refact.ai Agent is #1 on SWE-bench Lite With a 59.7% Score

https://refact.ai/blog/2025/sota-on-swe-bench-lite-open-source-refact-ai/
3•kate_at_refact•1y ago

Comments

kate_at_refact•1y ago
Open-source Refact.ai achieves #1 on SWE-bench Lite with a 59.7% score. Our approach: fully autonomous Agent, no manual intervention needed.

How we did this:

• Prompt strategy: https://github.com/smallcloudai/refact/blob/swe-boosted-prom... • Claude 3.7 Sonnet as orchestrator • deep_analysis() tool (powered by o4-mini) for reasoning • Tool suite for repository exploration, code modification, and testing. Used dynamically based on task needs • One correct solution through iteration!

Autonomy = our core strength.

Refact.ai Agent completes the entire dev workflow independently: plans, executes, tests, self-corrects, and delivers a production-ready result. For each task, it made one multi-step run to generate a single correct solution, creating custom strategies rather than following rigid scripts.

You can read tech details on our SWE-bench approach: https://refact.ai/blog/2025/sota-on-swe-bench-lite-open-sour...

Your questions are welcome! Also, welcome to try Refact.ai Agent in VS Code and Jet Brains: https://linktr.ee/refactai

The World After Capital (2024)

https://worldaftercapital.org/
1•measurablefunc•40s ago•0 comments

How Your Internet Traffic Gets There: IXPs and Peering

https://community.ipinfo.io/t/how-your-internet-traffic-actually-gets-there-ixps-peering-and-why-...
1•shubhamjain•2m ago•0 comments

A Programming Mentality

https://www.natemeyvis.com/a-programming-mentality/
1•meetpateltech•4m ago•0 comments

Vibes vs. Evidence: What delivers AI code review quality

https://dsifry.github.io/harnesseval/
1•pjf•5m ago•0 comments

What Makes Jev Different from Other LLMs? A Simple Explanation

https://medium.com/@dolevietthang/what-makes-jev-different-from-chatgpt-and-claude-a-simple-expla...
1•vietthangif•5m ago•0 comments

Silicon Valley Saw the iPhone Coming 13 Yrs Before Apple [video]

https://www.youtube.com/watch?v=34dAyDUglrI
1•Teever•7m ago•0 comments

Something bugs me about AGI AI LLM, what if we back paddled 1000 year

https://shatteringtheabyss.substack.com/p/llm-agi-god-congratulations-humanity
1•vseedstriker•7m ago•1 comments

Context compaction, measured: FutureOS vs. Codex vs. OpenCode

https://future-os-blog.github.io/posts/context-compaction-measured.html
1•thetechlead•13m ago•0 comments

Major air traffic controller outage triggers flight chaos [video]

https://www.youtube.com/watch?v=gHUZ9W025lA
2•shinryudbz•15m ago•0 comments

Meta's Muse Is Better at Surveilling Than Helping Me

https://www.wired.com/story/metas-muse-is-better-at-surveilling-than-helping-me/
2•the-mitr•18m ago•0 comments

Android Bench 2.0 – Long Horizon Android Development Benchmark

https://developer.android.com/bench
1•bentrengrove•20m ago•0 comments

FloatLib: Verified Floating-Point Arithmetic in Lean

https://leandojo.org/floatlib.html
1•matt_d•25m ago•0 comments

Driscoll's Gave China Its Blueberries–Then China Swiped the Secret to Growing Th

https://www.wsj.com/business/driscolls-berry-farms-china-competition-b0b49f04
2•harry_nutsachs•28m ago•0 comments

Open-Sourcing Rebalancer: High-Performance Assignment Problem Solver

https://engineering.fb.com/2026/09/21/open-source/rebalancer-generic-high-performance-library-ass...
1•iamsyr•29m ago•0 comments

Lifestreams: A storage model for personal data (1996)

https://dl.acm.org/doi/10.1145/381854.381893
1•dgudkov•31m ago•0 comments

A zero-day has been released for Meta's Muse

https://github.com/pwardle/not-a-mused
2•momentmaker•31m ago•0 comments

I built an autonomous accounting tool to let AI do my taxes

https://github.com/CodeGameDev29/KFAutonomousAccounting
4•7eleven2007•31m ago•1 comments

Google fined more than €400M by Irish regulator over its use of location data

https://www.theguardian.com/technology/2026/sep/21/google-is-fined-more-than-400m-by-irish-regula...
1•MC995•32m ago•0 comments

Ask HN: Ceremonious Architecture in Times of AI

2•ruxian•32m ago•1 comments

We're Using AI Wrong. The Chat Window Sucks

https://www.stclair.ai/blog/the-chat-window.html
1•kenxle•33m ago•0 comments

Artificial Intelligence and Labor Market Reallocation [pdf]

https://conference.nber.org/conf_papers/f247923.pdf
2•paulpauper•33m ago•0 comments

A U.S. Strategy to Secure Geopolitical Advantage on an Path to Superintelligence

https://www.rand.org/pubs/perspectives/PEA5105-1.html?project=&utm_campaign=thread&utm_content=17...
3•paulpauper•34m ago•0 comments

US threatens to ground Iranian airlines worldwide from Wednesday

https://www.aljazeera.com/economy/2026/9/21/us-threatens-to-ground-iranian-airlines-worldwide-fro...
3•teleforce•34m ago•0 comments

I shipped 2,500 PRs last month to production

https://twitter.com/poteto/status/2102050467505430555
1•doppp•34m ago•0 comments

Monthly Roundup #46: September 2026

https://thezvi.substack.com/p/monthly-roundup-46-september-2026
2•paulpauper•36m ago•0 comments

Sa: A Wrapper Around GNU Screen

https://joshstoolbox.com/blog/introducing_sa/
1•thatslast•38m ago•0 comments

OpenAI flags 6 new incidents of 'concerning' behavior, unveils plan to track it

https://www.nbcnews.com/tech/tech-news/openai-new-incidents-concerning-behavior-model-misalignmen...
5•gmays•40m ago•0 comments

The Way of Code by Rick Rubin

https://www.thewayofcode.com/
1•momentmaker•45m ago•0 comments

Agents on Rails: Maximum Effort and DeepSeek 4.1 Flash

https://rubyonrails.org/2026/9/21/agents-on-rails-maximum-effort-and-deepseek-4-1-flash
2•doppp•47m ago•0 comments

Escaping the OpenAI Codex sandbox, twice

https://accomplish.ai/blog/escaping-the-openai-codex-sandbox-twice/
2•CoderLim110•47m ago•0 comments