It also seems like the new Polonius/trait-solver work is pushing compiler performance toward a more interesting problem: doing expensive analysis only when it is actually needed.
4.57% mean wall-time reduction across 629 benchmarks in two months is pretty remarkable. Great progress.
Isn't this true for most optimizations, not just in compilers? My usual goto process for optimizing is "Find stuff we're doing that we don't have to do, re-evaluate what data structures we use and then re-evaluate what algorithms we use" basically, with minor changes depending on the results. Served me well so far, and haven't (intentionally) written any compilers.
What is the performance killer?
Generics/monomorphization and how iterators work results in a lot of compiler bytecode that has to be churned through. More bytecode = longer compilation. It increases the size of the (debug) binaries, the debuginfo in general, causes performance issues with debug binaries in some situations unless you bump the optimization level, causes more IO, etc.
Sometimes we really can have our cake and eat it too.
Now, is rustc slower than e.g. clang? by how much? why?
Those are different (and complicated) questions. It really depends on what you're compiling, but I'd say rustc can be 1-5x slower (maybe more at times?).
The reasons are many and varied, but in general rust compilation is slower because the compiler is doing way more things compared to C (monomorphization, complex trait resolution + type inference, borrow checker..)
[0]: Also I'd argue that "slow" without a concrete point of reference is a meaningless term in this context.
If you want something like Rust that offers guarantees and checks and cross-checks by the boatload, it adds up. Macros, monomorphization, implicit code generation with traits and all those other things add up too. And you can't always get O(n) or O(n log n) code to implement those checks. Maybe it can be sped up and maybe there's tricks here or there, but at the Pareto frontier, a language that has more checks will be slower to compile than one that has fewer.
And that's not a bad thing or a deficit in Rust, it's just the nature of the beast.
So the focus on Rust compile time is misplaced. It's not a big deal in terms of overall productivity.
In unoptimized builds often the linker is the bottleneck. Rust/Cargo can parallelize most of the build, generating tons of code and debug info, but then the poor linker has to consume all of it at once. The object/exe formats were designed in ancient times, so they're hard to build incrementally or in parallel (some linkers are trying).
Both are kind of outside the Rust compiler's influence. Macros can be almost arbitrarily complex: you pay for what you order. Codegen is LLVM, and that's a fixed choice. You can use Cranelift to get around it, but then you pay elsewhere.
Regardless, it's nice to see performance improvements in the compiler, even if you have to cooperate to benefit from them (e.g. by keeping your macros light).
torutofu•1h ago