I think the coding ability of AI agents may have really crossed a critical point recently. At least in some very hardcore software engineering tasks, it has begun to surpass the vast majority of human engineers and even entered fields that only a few experts could do well in the past. A month or two ago, I had Opus 5, Fable 5, and GPT-5.6/GPT-6 optimize the same mature open-source code repository. This is not a newly written project with low hanging fruits everywhere. It has been continuously maintained for over 8 years and used by a large number of real devices and software. Performance optimization itself has been one of the most important tasks of this project over the years. After repeated attempts by several of the strongest models at that time, they could probably only match or slightly surpass excellent human developers. But these days, I did it again with Opus 5.5. The changes are already very obvious. It is not about finding a simple optimization point, but rather conducting multiple rounds of profiling, analysis, implementation benchmark, Continue to search for the next bottleneck based on the results. In the end, on my test set, compared to the original implementation: The throughput can almost double at most; The average decrease in CPU computation is about 30%; The average energy consumption has decreased by about 30%; The average memory usage has decreased by about 30%. And these are still average returns, with greater improvements in certain specific scenarios. What really makes me interesting is not just 'AI has written the code a little faster'. But it is entering a field that relied heavily on the experience of a few experts in the past: Faced with a mature code repository that has been optimized by humans for eight years, it can still find new performance space, propose hypotheses, implement optimizations, validate with benchmarks, and continue iterating on its own. A month or two ago, I thought that AI agents had just caught up with the best engineers in this type of task. Now I'm starting to suspect that we may have already crossed that intersection.
--
Loading...