Insights for the Advancement of LLMs — Will Semiconductor Demand Truly Explode in the AI Era?

The Aesthetics of Subtraction II — will semiconductor demand in the AI era truly explode?

An Seungwon · Wonbrand CEO · July 27, 2026


Preface

With the dawn of the AI era, a blind faith in exploding semiconductor demand and faster processing speeds has dominated the market. For the past few years, Large Language Model (LLM) research has pursued only the aesthetics of addition: 'more and faster'. However, more computation and faster inference are rapidly approaching their physical and efficient tipping points.

Technological advancement is not a linear expansion; true innovation stems from an accurate recognition of limitations and a structural transition. What we need now is not a blind race for semiconductors, but a fundamental reflection on what to subtract and how to become more accurate. As an extension of my previous proposal, 'The Aesthetics of Subtraction', I present four insights to dispel the current illusions of the AI industry and point out the concrete direction the next generation of AI must take.

Proposal 1. The Illusion of Exploding Semiconductor Demand and Ecosystem Oligopoly

It is self-evident that semiconductors are the core infrastructure of the AI industry. However, interpreting the phenomenon of countless companies pre-purchasing years' worth of chips as a simple 'demand explosion' is a superficial analysis. This is driven by cold calculation. It functions largely as 'strategic hoarding'—the logic being that securing chips deprives competitors of them, thereby providing a competitive edge.

While many companies are currently challenging the AI space, the market will eventually consolidate around a very small number of top-tier companies like Anthropic and OpenAI. Once competitors decrease and an oligopoly forms, the blind race to secure hyper-scale infrastructure will subside, naturally reaching an equilibrium of appropriate inventory and demand. We must critically rethink the illusion that semiconductor demand will curve upward infinitely.

Proposal 2. The Limits of Probability-Based Models and the Potential Collapse of the Semiconductor Paradigm

Unless there is a revolutionary technological discovery that changes human destiny, it is nearly impossible for latecomers to close the technological gap with current leading companies overnight. However, this premise is only valid as long as the current architecture is maintained.

We must face the inherent limitations of the current AI framework (autoregressive models) based on the 'probability' of the next word. Pouring infinite data to calculate statistical probabilities will eventually hit the wall of logical reasoning. Perhaps soon, a revolutionary discovery will completely transition this Transformer-based framework into a non-autoregressive or an entirely different, highly energy-efficient architecture. The moment this paradigm shift occurs, the massive semiconductor (GPU) infrastructure optimized for probability-based matrix multiplication—and the corporate war to secure these chips—could become obsolete overnight. The technological gap built solely on the sheer volume of chips might just be a castle made of sand, ready to be invalidated in an instant.

Proposal 3. Abandoning Blind Speed in Pursuit of "Appropriately Slow" Accuracy

Most AI users, including myself, do not just want faster processing when using current AI technologies; we want an agent that 'accurately understands' our instructions, even if it takes a bit more time. It is completely meaningless for an AI to spit out results rapidly if it has misunderstood the prompt.

This is clearly felt when using current LLMs. When instructing an AI to code, if it uses the trick of reading only a snippet to process it quickly, the result is always poor. Conversely, even if processed moderately slowly, an approach that reads the entire code and fully understands the overall context (System 2 Thinking) always yields excellent results. User instructions are now becoming more complex, spanning videos, photos, and design, far beyond simple text. Accordingly, 'accuracy'—willingly consuming inference-time compute to perfectly understand and execute instructions—will become the most important value of AI, rather than meaningless 'speed'.

Proposal 4. Reaching the Tipping Point and AGI Engineering: The Aesthetics of Subtraction

The race toward more data, more massive parameters, and faster inference has already reached the tipping point of efficiency. The era of one-dimensional scaling, blindly increasing model size, is setting. The core of the truly impending AI era is not size expansion. It lies in the 'Engineering' of making the model as accurate and smart as possible within limited resources, and the bold 'subtraction' of the unnecessary.

Personally, I am conducting research to realize AGI (Artificial General Intelligence) by layering sophisticated engineering architectures on top of current LLMs. In this process, the three core pillars I consider most critical for reaching AGI are: First, 'Cross-domain thought transfer', which smoothly applies knowledge and logic learned in one domain to problem-solving in a completely different domain. Second, 'Resistance to forgetting', which ensures the core self and important past contexts are not lost amidst a constant influx of new information. Third, a 'Unique Persona', which allows the model to function as an independent entity with consistent values and response systems, moving beyond a simple chatbot that resets every session.

The art of subtraction, which compresses massive data and leaves only core reasoning abilities, combined with the aforementioned meta-engineering, is the only solution for the next generation of AI to break through its limits and be reborn as true intelligence.

Closing

The curve of technological advancement has always led to convergence toward the essence after a period of blind expansion. AI is no exception. The exhaustive competition of blindly hoarding semiconductors and increasing model size is nearing its end. Not a machine that skims and answers quickly, but an AI that takes the time to read code to the end and understand it accurately. An AI that does not obsess over size, maintains a unique persona, and knows how to smartly subtract from itself.

This is the true face of the next-generation artificial intelligence that the 'Aesthetics of Subtraction' will complete. It is time for us to break free from the myth of speed and expansion and ponder deeply once again on the essence of technology.


References & Further Reading

  1. On the Limits of Autoregressive Models: LeCun, Y. (2024). Objective-Driven AI and the Fallacy of Scaling Laws.
  2. On System 2 Thinking in AI: OpenAI (2024). Learning to Reason with LLMs (Project o1).
  3. On Semiconductor Market Dynamics: Thompson, N. C., et al. (2025). The Compute Divide: Analyzing GPU Hoarding and Utilization Rates in Big Tech.
  4. On AGI and Engineering: Marcus, G. (2025). Towards AGI: Cognitive Architecture, Persona, and Cross-Domain Transfer.

An Seungwon / Wonbrand / https://wonbrand.co.kr