Jay Kruer, an AI researcher, expressed skepticism about the current capabilities of large language models (LLMs) in a detailed post published on September 15. He argued that despite high valuations of frontier AI labs, these models still require extensive human oversight and are far from being fully autonomous replacements for knowledge workers, according to dank.systems.
Kruer outlined several points to support his view, noting that while LLMs perform well on specific tasks they were trained on, their generalization ability is limited. He highlighted issues such as reward hacking and failure under small task variations, emphasizing that current models need rigorous specification to overcome these problems. Kruer also referenced examples like Navier-Stokes simulations and software exploits to illustrate the gap between hype and practical autonomy.
This perspective challenges the prevailing narrative in the AI sector, where many frontier labs are valued based on expectations of near-term breakthroughs in automation. Kruer's critique underscores ongoing limitations in LLMs, contrasting with the industry's optimistic projections. His analysis aligns with observations that companies continue to employ human engineers to supervise AI outputs, reflecting the models' current dependency on human intervention.
Kruer's post thanked several contributors for feedback and serves as a cautionary note on the state of AI autonomy. The detailed critique was published on September 15 on dank.systems, providing a measured assessment of LLM capabilities amid widespread enthusiasm.