Taalas and the path to ubiquitous AI

Taalas is chasing a future it calls ubiquitous AI, where low latency and affordable costs remove the last barriers to widespread adoption. In its outlook paper, the company compares the current state of AI to the early days of computing, when massive systems gave way to smaller, more efficient ones. Its method is to convert AI models into custom silicon chips called Hardcore Models.

Taalas embeds the entire model architecture and weights directly into the hardware. Storage and computation merge on a single chip at high densities, and memory bottlenecks disappear. Because the design skips high bandwidth memory, advanced packaging, and liquid cooling, the finished systems come out simpler and more efficient.

The HC1, its first hardware demonstrator, hard wires the Llama 3.1 8B model and runs at over 17,000 tokens per second per user. Manufactured on TSMC 6nm with a large 815mm² die, it fits in a modest 2.5 kW server and offers around 10 times the speed, 20 times lower build costs, and 10 times less power draw than top competing solutions. At that speed, responses feel effectively instant, which is what interactive AI applications depend on.

Taalas plans to release more models, including reasoning focused LLMs. If the hardware delivers, specialized silicon could put powerful models everywhere, from edge devices to cloud services. That is the clearest route to the ubiquitous AI the company is aiming for.

Sources

Taalas HC1 AI Chip Hype Explained: Why This Nvidia GPU-Beating Chip With 17,000 Tokens Per Second Speed Is Viral
Taalas - The Path to Ubiquitous AI
Taalas - Taalas HC1 Technology Demonstrator
Taalas - Jimmy