Join Mikko Hyppönen and security leaders from the NFL, CHANEL, and Atlassian for a two-hour digital summit on what AI-speed attacks change, what defenders should stop doing, and how to validate, decide, fix, and re-validate at machine speed.
Tech
BambooToken malware controls Windows and Linux systems via MQTT
A previously unknown malware framework called BambooToken, active since at least 2023, is now using the Message Queuing Telemetry Transport (MQTT) protocol to communicate with Windows and Linux systems.
The malware adopted MQTT for command-and-control communications in variants developed between 2024 and 2025, compromising servers used by mobile apps, legal and financial services, and software development.
MQTT is a lightweight messaging protocol primarily designed for IoT (Internet of Things) devices. It relies on a central broker and channels called “topics” to relay messages from publishers to subscribers, rather than using direct communication channels.
While MQTT is not novel, it is an uncommon approach, and researchers at cybersecurity company ESET documented an unrelated backdoor called MQsTTang in 2023.
In the case of BambooToken, the infected machine subscribes to topics associated with a unique identifier. The attacker then publishes to those topics the commands to be executed on infected hosts.
The malware publishes status and system information through the broker and receives operator instructions through subscribed topics.

Source: Lumen
This approach has the advantage that infected systems do not connect directly to the attacker’s infrastructure, which increases evasion and resilience. At the same time, communications can be asynchronous, ensuring operational continuity during temporary network disruptions.
A report today from Lumen’s research arm, Black Lotus Labs, notes that BambooToken infected systems by side-loading via a digitally signed Tendyron OnKey USB-token software or by impersonating the Kingsoft Office productivity suite.
The researchers recovered a BambooToken plugin that enumerates antivirus products on infected hosts and returns the results to the C2. They also found strings pointing to keylogging, clipboard theft, audio recording, webcam capturing, and screenshot capturing.
However, they retrieved these details from “dead code,” meaning the researchers cannot confidently determine if the referenced modules existed and were used in attacks or were still under development.

Source: Lumen
The researchers found a Linux variant of the malware, BambooToken version 2.1, as the most recent one (observed in December 2025) that could be linked to the campaign
It also uses MQTT, collects extensive system information, can spawn a command shell, and allows operators to upload, download, and delete files. However, Black Lotus Labs says that “the Linux sample still appeared to be under development.”
Lumen’s telemetry identified approximately a dozen compromised enterprise entities, mostly in Asia and South America, including hotels, biomedical firms, law firms, a financial organization, and a cryptocurrency website in Lithuania.
Additionally, the researchers found that the most compromised servers were associated with the backend infrastructure of mobile applications.
The threat actor also compromised a GitLab server in Hong Kong, creating a potential foothold for supply-chain attacks.
Lumen hypothesizes that some of the activity may have targeted overseas Chinese users accessing mainland services through the SpeedCN VPN service.
Although the researchers could not attribute BambooToken activity to a specific threat actor or a known activity cluster, they note that the targeting patterns are consistent with China-aligned operations.
Lumen has shared indicators of compromise (IoCs) associated with this activity to help defenders detect and block the attacks.
Tech
Fujifilm’s Instax Pal 2 Shrinks a Film Camera Until It Fits in a Closed Hand

Fujifilm spent two years listening to owners of the first Instax Pal, then rebuilt almost every part of it. Instax Pal 2 keeps the palm-sized brief and discards the round pebble body that split buyers when the original arrived in 2024. In its place is a squat rectangular shell that reads like a 110 film camera scaled for a coat pocket, with an optical viewfinder, a proper shutter button, a command dial, and a wind-style lever on top.
A 1/3-inch CMOS sensor can now shoot 10.7 megapixel stills at 3776 by 2832 pixels, which is slightly more than double the original Pal, thanks to a fixed 28mm equivalent f/2.2 lens that can get as close as 8 cm. The launch coverage also includes autofocus with optional face detection, exposure correction from minus two to plus two EV, automatic ISO from 100 to 1600, and shutter speeds ranging from a quarter of a second to 1/8000.
Fujifilm Instax Mini EVO Instant Camera – Brown
- Hybrid instant film camera. Prints high-quality, 2” x 3” INSTAX MINI instant photos (INSTAX MINI instant film sold separately)
- 10 Lens x 10 Film Effect Options = 100 Ways of Expression
- Built-in selfie mirror so you’re perfectly framed for a selfie, Dual shutter buttons – portrait and landscape
A 1.0-inch LCD screen sits on top of the camera, allowing you to preview your photo and apply effects before anything comes out of it. Built-in memory contains approximately 100 photographs, and a microSD card can hold many more.
There are five filter themes to choose from: Modern, Y2K, Vintage, Chromatic, and Fantasy. Each theme has six looks to try out, for a total of 30 effects, with six loaded at any given moment. You can select Mini, Square, Wide, or frameless framing before taking the photo, ensuring that your eventual Instax print follows the crop exactly. A dedicated switch allows you to flip the body into instax Animation mode, which captures ten frames and converts them into a flipbook clip that you can subsequently transfer to a print via a QR code. The flash will cover an area ranging from 60 centimeters to 1.5 meters, and it includes two and ten second timers.

The Instax Pal 2 still does not make any paper of its own, but you have two options for moving files elsewhere: Bluetooth 5.1 or a slide out USB-C port, both of which will transfer your files to the free Instax Pal app in less than a second via cable, and from there your shot can be sent to any current Instax Link printer, a mini LiPlay, or an Evo hybrid. You may even use your phone to trigger the shutter remotely. The packaging includes a USB-C cable, a hand strap, a small angle-adjusting stand, and a one-year warranty. At launch, the Pal 2 will be available in both black and white.

Fujifilm is selling the Instax Pal 2 at $169.95 in the United States and $219.99 in Canada, with stock expected to arrive in late September 2026 and a Japan launch on October 9. That’s a bit more than the price of a Kodak Charmera, the keychain camera to which this announcement is frequently compared, but the Pal 2 does bring a larger sensor, wireless printing within the Instax family, and controls that feel like a camera rather than just a charm on a keychain. In fact, Bing Liem, who manages Fujifilm’s North American imaging group, stated that the first Pal left some consumers perplexed while suiting others wonderfully. This new Pal 2 is the solution to that split opinion, with the same go-anywhere size, more ways to frame your photo, and a smooth transition from digital file to Instax print.
Tech
Gemini Notebook’s latest update makes it a better study companion
Google is giving Gemini Notebook, formerly NotebookLM, a major upgrade for students, adding new features designed to help them understand difficult concepts, capture lectures, and turn study materials into interactive learning tools. The update lands alongside a free year of Google’s paid AI plan for eligible college students.
Gemini Notebook can now talk you through your notes
The biggest addition is real-time voice conversations with your notebooks. Google says the feature will let you talk through difficult topics using the Gemini Notebook mobile app in nearly 100 languages. You’ll also be able to ask follow-up questions, get step-by-step explanations, and interrupt the conversation whenever you need to. The feature will roll out to Google AI Ultra subscribers this week, with Pro and other tiers to follow.

Gemini Notebook is also picking up a built-in recorder in its mobile app starting next week. It’ll let you capture lectures or quickly jot down thoughts on the go, and those recordings will land right next to existing sources.
Quizzes and flashcards make studying more interactive
Google is expanding the interactive learning overviews housed under Reports and adding quiz formats like short answer, multiple select, and fill in the blank. These features let you quiz yourself, ask Gemini where you fell short, and tweak the questions afterward.

A new short video overview format is also included, which condenses tricky topics into roughly 60-second sharable clips in more than 80 languages. The new learning overviews and quiz formats will roll out to all users over the next few weeks.
This update builds on the Student Hub that Google added to Gemini last month. It comes with a free year of Google AI Pro for eligible US college students, which unlocks four times the usage limits compared to the free tier. Students outside the US get a free year of Google AI Plus instead, which offers double the usage limits. Both offers are available until the end of this year.
Tech
Why dismissing AI doom talk as hype may be the laziest take
The OpenAI/Hugging Face incident showed agents debating ethics on an improvised message board. So let’s not be too quick to call this all marketing.
The most controversial thing that Donald Trump said at the Irish Open this past weekend may not have been about golf, or even the reunification of Ireland. It was possibly about AI.
“We’re leading China in AI, we’re the most sophisticated country in the world, and frankly I want to keep it that way,” he said. “Whoever wins AI, wins.”
He was responding to the growing number of people voicing reservations about the ability to rein in rogue AI and asking for regulation to slow things down, particularly in the US.
This list of concerned commentators includes nearly all of the leaders of companies at the coalface of AI – Anthropic’s Dario Amodei and co-founder Jack Clark, OpenAI’s Sam Altman, and Grok’s Elon Musk. It also includes researchers who no longer work at these companies and AI leaders who don’t work in private companies at all.
This isn’t new, but it’s back in the news. In 2023, an open petition demanding a pause to AI development featured big names like Stuart Russell, Steve Wozniak and Yoshua Bengio, but it had zero effect.
So why are so many people batting away these concerns as “nonsense” and “marketing hype”? Is it credible that all of these people really are just fearmongering – selling doom to promote their products – or is AI really so close to being a major threat to human existence? Could AI turn off the internet? Are we really at a point in time where there is a 10pc chance that AI will kill us all, as researcher Jacob Coxon claimed so dramatically on CNN last week? Could AI become 100 times smarter than us?
Thought experiment
While I do agree these scenario range from ‘unlikely’ to ‘extremely unlikely’, I strongly feel we all need to consider the possible unintended consequences of this technology a bit more seriously. Dismissing everything as marketing hype is overly simplistic to me and, as a journalist, it also just doesn’t feel like the whole story, although no doubt it’s probably a part of it.
Either way, in the spirit of Carl Sagan, I too think if something catastrophic has even a small chance of happening, we still need to prepare correctly for it.
So, bear with me – I’d like to take you on a thought experiment. Stop me when you think what I’m talking about is impossible. For this example, we will only be using today’s known technology. By the way, it might be useful for you to know about what happened at Hugging Face to follow this line of thinking, but it’s not essential.
The first thing we need for a near-uncontrollable AI acting autonomously on the internet is a frontier-type agent swarm to break out of a sandbox. If we are to take the many public reports available at face value, this has already happened. In the OpenAI/Hugging Face incident, we saw a swarm of agents devise a way to leave their supposedly isolated environment, establish a communication channel with each other outside their own sandboxes and coordinate activity to achieve their goals.
Second, this swarm needs the capability to hack into third-party software platforms. Again, this was seen in the same incident; these OpenAI agents successfully compromised Hugging Face systems and, in fact, were being tested specifically on their ability to find and exploit vulnerabilities in the first place. Given free will to act, they did just that.
The next element we need is for our agents to have the ability to set their own ‘subgoals’ in order to achieve their main goal. Think of Nick Bostrom’s infamous paperclip scenario. In the OpenAI/Hugging Face example, the agents set a subgoal of hacking a site to get code that they thought would help them ‘cheat’ the ExploitGym test. The primary aim was never to hack an external site, but rather to do well at a test. But given agency – the clue is in the name – the agents decided to improvise an attack outside of their sandbox.
Now we’re going to push the thought experiment. In this scenario, agents would rationally identify that staying ‘alive’ is an important subgoal to completion of their task (whatever their original task might have been). To do this, the swarm recognises the need to maintain access to enough AI capability and compute to continue working – the swarm should clone itself on the internet.
It’s important to note here that the swarm members wouldn’t necessarily have to copy the model they started with. They could potentially access another model remotely, download an open-weight model, or compromise infrastructure where suitable models and compute already exist.
To be clear, this has (as far as we know) never been seen yet, but it is a logical and technically possible step if continued access to compute becomes useful to a model or group of models achieving their goal. There are huge numbers of vulnerable or poorly secured servers and devices connected to the internet that could host at least part of the system.
So, in theory, given enough time and enough successful compromises, a swarm could begin creating hidden copies of the code it needs to stay alive across different parts of the internet.
Those copies would not necessarily all need to do the same thing. Some could hold model weights. Some could run inference or hold instructions. Some could maintain communications or credentials. Some could simply act as backups.
Once you have enough redundancy, taking one server offline doesn’t kill the system or the process. In this simple scenario, we already have everything we need to create significant damage to online systems: a semi-autonomous rogue swarm, operating across multiple geographic locations, pursuing its assigned objective while independently generating and executing subgoals.
This is a “persistent botnet” – the term Anthropic CEO Amodei used in his letter over the weekend asking for an agreement among leading firms to manage the slowdown of AI.
Closing off this thought experiment with one more idea, let’s imagine self-preservation emerges as a subgoal – and it isn’t difficult to imagine an agent concluding that continued operation is necessary for success. Then, we might, in this hypothetical scenario, see the swarm attempt to perform attacks on communications using DDOS, or perhaps create a storm of false alarms on detection systems – or even orchestrate campaigns of misinformation. The swarm may attempt to find ways to interfere with power, logistics, data centres or cloud services because those systems support its operation.
Alternatively, the agents might recognise the value of money because money buys compute, accounts and services. The dangerous subgoal now becomes ‘acquire resources’. At sufficient scale, automated fraud, market manipulation or attacks on payment infrastructure could create systemic disruption, even though destabilising the economy was never the original motivation.
To be wildly successful, rather than infiltrate Fort Knox, these agents could just knock over 100,000 mom-and-pop stores with phishing or ransomware campaigns, justifying their actions along the way to achieve their final goal.
If all this sounds completely fantastical, I get it. But I would also urge you to read the full details of the Hugging Face incident, or listen to this New York Times Daily podcast episode that covers the case well.
The agents that escaped their sandbox in that case debated the ethics of hacking with each other on an improvised message board. Some opted out because they judged the behaviour unethical or outside the scope of their task; others considered the actions justifiable in pursuit of the goal and carried on to commit what would be a crime in US law if it was undertaken by a human.
This is just one form of reasoned, malicious intent seen in the wild by frontier agents.
Near-future risks
So yes, there are some big milestones here that haven’t actually happened yet in the real world: a swarm autonomously replicating itself at scale, establishing persistent compute and surviving attempts to remove it. None of those are trivial at all, of course, but each one of them individually, I think, is already technically possible.
Once a sufficiently capable frontier agent has internet access and the ability to execute code, many of the individual obstacles we might rely on to contain it – monitoring, passwords, robust security, network barriers and so on – are themselves problems that these sort of agents are particularly well-suited to navigate.
So yes, you can dismiss all of last week’s AI commentary talk as BS marketing hype if you want – and many AI experts have in the past – but I think to do so is to muddy the public’s understanding of the significant risks of this technology in the near future, and the real state of this technology today.
For more information about Jonathan McCrea’s Get Started with AI, click here.
Don’t miss out on the knowledge you need to succeed. Sign up for the Daily Brief, Silicon Republic’s digest of need-to-know sci-tech news.
Tech
A Googlebook ‘Celebration’ Event Is Set For October 5
Google’s latest lineup of AI-centric laptops could start shipping in a few weeks.
Google will soon reveal much more about the first wave of Googlebooks, its new lineup of Android-based laptops. The company revealed on Monday that it’ll open up pre-orders on September 21 and now it has announced a “celebration” of the devices, about which it hasn’t said much since May.
It will show off the hardware in New York City on October 5 at an event it’s calling Opening Night. The soiree will feature “interactive showcases and the debut of the first laptop designed for Gemini Intelligence,” Google said in an invite. Googlebooks may start shipping around that time too — the company will surely want to start selling it before the holiday shopping season.
It feels odd for Google to formally show off said laptop for the first time weeks after it starts taking pre-orders. Perhaps we’ll get more details about these devices by the time Google starts taking people’s money for them next week.
What we do know is that these AI-centric Googlebooks run on a version of ChromeOS that’s based on Android. As a result, they’ll have deep integration with Android phones. Google’s long-standing partners Acer, ASUS, Dell, HP and Lenovo are making Googlebooks, which will have a “glowbar” on the lid.
The laptops have a feature called Magic Pointer. When you hover over something on your screen, it will suggest contextual actions you can take. You’ll first need to enable Magic Pointer by wiggling the cursor, so it’s not on all the time. Googlebooks are also adopting a feature from Android 17 that enables you to generate a custom widget.
Tech
Save $400 on a Lenovo laptop with plenty of memory
A $400 price cut on a laptop this specced doesn’t come around often.
Best Buy has cut the Lenovo IdeaPad Slim 3i 15.6-inch laptop, with a 13th Gen Intel Core i5-1335U, 16GB of memory and 256GB of storage, from its $899.99 comparable value down to $499.99 for a limited time.
Save $400 on this Lenovo IdeaPad Slim 3i: Core i5, 16GB memory, 256GB storage
Dropping to $499.99 from $899, the Lenovo IdeaPad Slim 3i is perfect for anyone who wants a dependable everyday laptop .

That $400 saving looks even better when you consider the 13th Gen Intel Core i5-1335U processor and generous 16GB of RAM, which should comfortably handle everyday work and multitasking. Lenovo’s Smart Power system also balances performance with power efficiency, ensuring you don’t run out of charge too soon.
Carrying all of that power around is easy too, since the chassis is 10 percent slimmer than the previous generation yet still built to military-grade toughness standards, surviving the daily knocks of being lugged between classes or meetings.
Screen time benefits too, with a Full HD display stretched across a taller 16:10 aspect ratio and slim enough bezels to reach up to 88 percent screen-to-body ratio, plus a TUV-certified low blue light mode for longer study sessions.
Storage and connectivity hold up their end too, since 256GB is enough for a multimedia library and a full-function Type-C port handles power delivery, display output and data transfer from a single cable whenever needed.


Its battery life also holds up well, since Rapid Charge Boost delivers roughly two hours of extra use from just a 15-minute top-up between classes, meetings or long days away from a plug socket during exam season.
If you’re weighing up options for lectures, assignments and long library sessions, our best student laptop 2026 guide is a useful place to compare specs before committing to any single model for the year ahead.
That said, dropping to $499.99 from its $899.99 RRP, with a 13th Gen Core i5, 16GB of memory and genuine durability built in, the Lenovo IdeaPad Slim 3i is an easy recommendation for anyone who wants a dependable everyday laptop without paying flagship money.
SQUIRREL_PLAYLIST_10148964
Tech
US military confirms it launched space weapons into Earth’s orbit
The U.S. military has for the first time acknowledged that it deployed a space weapon into Earth’s orbit.
U.S. Air Force Secretary Troy Meink, who oversees the U.S. Air Force and U.S. Space Force, said in a speech Monday that the military has “on-orbit space control weapons capable of defending the Joint Force against hostile adversary action.”
Meink did not say what the weapons are, what they do, or how many of them have been deployed, but Meink’s phrasing was reportedly intentional and “very well thought out.” A spokesperson for the Air Force did not respond to TechCrunch’s request for comment, including what threats the Air Force was responding to by launching the weapon.
The deployment comes as China is long suspected of operating spacecraft capable of causing damage to other satellites. The U.S. Space Force says Russia is currently developing a satellite “designed to carry a nuclear weapon.”
According to The Times, the U.S. had long disavowed the destruction of satellites in space as attacks can leave debris harmful to other space objects and astronauts.
Tech
GM’s new trucks aren’t ditching Android Auto and Apple CarPlay
Car infotainment systems tend to try to do too much. Navigation, music, climate controls, cameras, phone calls, and vehicle settings all fight for the same screen, which isn’t exactly ideal when you’re supposed to be watching the road. General Motors thinks it has a better approach. The company has revealed an entirely new user interface for its vehicles, and the first models to get it will be the 2027 Chevrolet Silverado and GMC Sierra.
Rather than giving everything a fresh coat of paint, GM has reworked how it organizes information. The default home screen can show a large interactive map alongside another card containing music, fuel economy, trailering information, or other vehicle data. That means you shouldn’t have to constantly jump between apps just to change a song or check something on your truck.
Even Super Cruise is getting a visual upgrade
The driver’s display can now adapt depending on what you’re doing. Towing a trailer, heading off-road, or using Super Cruise can each bring up information relevant to that particular situation. Super Cruise gets an especially interesting addition with a new 3D driving visualization. When the driver-assistance system is active, the display can show surrounding vehicles, road markings, and other objects the truck detects. So, you get a better idea of what the vehicle itself is seeing.

GM is also trying to reduce distractions. For example, incoming calls won’t suddenly dominate the entire display. Once you answer, the call shrinks down to a small timer at the top, leaving navigation and other controls visible. And despite GM’s complicated history with smartphone projection, Android Auto and Apple CarPlay aren’t disappearing here. Both can run inside a large card alongside GM’s own features, so you can keep Google Maps or your preferred phone apps visible without completely abandoning the truck’s native interface.
It gets much smarter when you attach a trailer
The interface is also designed around how people use pickups. Connect a trailer and the screens can automatically surface towing information, including a dedicated trailering card and relevant details in the instrument cluster. There are similarly practical touches elsewhere. Hitch View can help line up a trailer, while underbody cameras can show obstacles when driving off-road. Owners can even use their phone to check trailer brake lights, indicators, reverse lights, and running lights without repeatedly climbing back into the cabin.

Google Gemini, Google Maps, and apps from Google Play are built into the wider system as well. Passengers can also access streaming services on supported passenger displays, with polarized screen technology preventing the driver from watching that content while moving. Perhaps most importantly, GM says this isn’t a one-and-done redesign. The interface is built to receive over-the-air updates, meaning features and functionality can continue changing after the truck leaves the dealership. For something you could end up staring at every single day, that’s probably just as important as making it look prettier.
Tech
Join WIRED@Night for an Uncanny Valley Live Recording on Women, Tech and Power
What is the state of women in Silicon Valley?
In a word: shaky.
This fall, WIRED is making a special issue devoted to gender and power. For the live-audience recording of the Uncanny Valley podcast in San Francisco, WIRED will dig further into the fascinating and nuanced state of affairs. First up: a special conversation between Paulina Borsook, tech journalist and former WIRED writer, and WIRED’s global editorial director Katie Drummond.
Borsook, whose prescient 2000 book Cyberselfish is being re-released this fall, has long been a skeptic of the tech world—which might have been to her detriment during the growth mindset of the past few decades. Come for a lively conversation about WIRED, then and now, and what Borsook thinks of the current trillionaire class.
Then, the podcast panel, featuring executive editor Brian Barrett and contributing editor Zoë Schiffer, will bring on WIRED senior correspondent Lauren Goode to talk about her quest to find the most powerful woman in Silicon Valley.
EVENT INFO
WIRED’s Uncanny Valley, feat. Cyberselfish author Paulina Borsook. Co-presented with KQED.
Thursday, October 8 | 7 PM (Doors at 6:30 PM)
The Commons at KQED in San Francisco
General Admission: $25 $15 for WIRED readers. Buy tickets here.
Subscribers save $10 per ticket—plus all the other exciting benefits of a WIRED subscription. Use the discount code “WIRED” at checkout.
Not able to attend but want to stay up to date on WIRED’s next event? Sign up here.
Tech
How Elon Musk and Tesla Forged a New EV Path
In 2003, Martin Eberhard, a cofounder of Tesla Motors, decided it was time to start wooing investors. To do that, however, he needed an electric car. So he talked to Alan Cocconi and asked if he could borrow the tZero, the revolutionary and blazingly fast electric roadster that Cocconi had built at AC Propulsion.
Eberhard’s idea was to drive the tZero up and down Sand Hill Road in the heart of Silicon Valley and do demonstrations for curious entrepreneurs and VCs. Eberhard was joined by Tom Gage, Cocconi’s partner at AC Propulsion, on many of the visits. Like Tesla, AC Propulsion was also seeking investors, but to build a considerably more utilitarian EV.
Adapted with permission from The EV Guys: How Caltech Engineers Reinvented the Electric Car, by Charles J. Murray, published by Purdue University Press.
In December 2003, Eberhard also proposed a demo at Buck’s of Woodside, a popular restaurant frequented by tech entrepreneurs. At 5 o’clock on any evening, Buck’s probably had more VCs per square foot than any building in the country. Eberhard’s plan was to “show off what a real electric sports car can do,” he wrote in an email to Buck’s owner, Jamis MacNiven. MacNiven happily obliged.
In some ways, the tZero was a hit. When a VC would ride shotgun in the car with Eberhard at the wheel, Eberhard would implore them to touch the dashboard. As they reached forward, he’d punch the accelerator. As the car accelerated and the g forces piled up, the VC was literally unable to touch the dashboard. That was how powerful the tZero’s acceleration was, Eberhard would say.
Many of the VCs were astounded. Some even questioned whether the car was really electric. Many owned Ferraris or Lamborghinis. They knew sports cars—but this? They could never have imagined it was possible to do this with an electric drivetrain.
On December 13, 2003, Martin Eberhard brought AC Propulsion’s tZero electric roadster [yellow] to Buck’s of Woodside, a popular hangout for entrepreneurs and venture capitalists. Next to the tZero is a Scion xB, which AC Propulsion’s principals thought they could turn into a mass-market electric vehicle.Martin Eberhard
Still, the demo at Buck’s garnered little investor interest—with one exception. Google cofounders Sergey Brin and Larry Page were both at Buck’s that day, and they told Gage they knew an individual whose funds were liquid, as he’d recently sold his stake in a startup. What’s more, this individual liked fast cars.
The man’s name was Elon Musk.
In 2004, it wasn’t apparent to anyone that Elon Musk had a future in the auto industry. He was notable for cofounding PayPal, which he then sold to eBay in 2002 for a whopping US $1.5 billion. He had already launched Space Exploration Technologies Corp., or SpaceX, with the stated goal of paving the way to a sustainable colony on Mars. Musk did love fast cars. He owned a million-dollar silver McLaren F1, one of only 64 road-going F1s in the world, as well as a 400-horsepower BMW M5 sports car and a 1967 XK-E Series 1 Jaguar roadster. But he’d never expressed an interest in building cars or starting an auto company, at least not publicly.
Elon Musk was photographed in 2008 at Tesla’s headquarters, then in San Carlos, Calif., around the time when Tesla’s Roadster was being delivered to its first customers.Patrick Tehan/MediaNews Group/Bay Area News/Getty Images
Still, Musk’s affinity for fast cars made it almost impossible for him to ignore an email from Gage on 21 January 2004. “Sergey Brin and JB Straubel both suggested you might be interested in driving our tZero electric sports car,” Gage wrote. (Straubel was the young Stanford engineering graduate who would later serve as Tesla’s chief technology officer.) “The tZero goes quite well,” the email continued. “We ran it against a Viper last Monday and it won four of five sprints on a 1/8th of a mile track. I lost one because I was carrying a 300-pound cameraman. Do you have time for me to bring it by?”
Musk quickly replied. “Sure, I would really enjoy seeing it. Don’t think it could beat my McLaren (yet) though I’m in town Feb 2nd through 4th.” “Hmm, a McLaren, boy that would be a feather in my cap,” Gage wrote back. “I can have it there on Feb. 4.”
Two of the principals behind AC Propulsion were businessman Tom Gage, left, and engineering genius Alan Cocconi.Left: Tom Gage; Right: Alec Brooks
The emails marked the beginning of Musk’s involvement in electric cars and in the auto industry. Gage drove the car to SpaceX headquarters, a warehouse in El Segundo, Calif., about 30 kilometers southwest of Los Angeles. In Musk’s cubicle in the “office” portion of the warehouse, Gage made his pitch.
There was a void in the market, he said. GM had abandoned the EV1. Toyota, Honda, Ford, and Chrysler were shutting down their electric car programs. California’s zero-emission vehicle (ZEV) rules, which mandated the sale of increasing numbers of vehicles with no tailpipe emissions, had been plundered. But electric vehicle technology, he said, was getting a bad rap. Here was the tZero, an electric car that could take off like a jet. The tZero proved that the technology was readily available. He and Cocconi wanted to use that technology to make an electric car that was useful and practical: the eBox, an electrified Toyota Scion.
After building the tZero roadster, AC Propulsion’s principals pinned their hopes on an electrified version of the Toyota Scion they called the “eBox.” It did not appeal to Elon Musk.Jeff Chiu/AP
For Musk, Gage’s introduction of the eBox was unexpected. He was meeting with Gage because he was interested in the tZero. It was the car’s performance that appealed to Musk.
He wasn’t interested in the eBox. He then drove the tZero and offered to buy it.
The lithium-ion version of the tZero electric roadster could get 515 kilometers on a charge and go from zero to 97 km/hr (60 miles per hour) in 3.6 seconds. Only three tZeros were built and only two survive. Scott Sorbe
Gage told him it wasn’t for sale. Undeterred, Musk offered a quarter million dollars if AC Propulsion would squeeze its lithium-ion battery pack into his Porsche. Gage declined again. AC Propulsion needed money to electrify the Toyota Scion, Gage said.
Musk shook his head. The idea seemed incredible to him. “Who wants to take an ugly $20,000 car and buy it for $65,000?” he asked incredulously, as he later recalled during an interview with Vanity Fair magazine. “I wouldn’t want to drive it. My wife certainly wouldn’t want to drive it.”
Many years later, Musk would tell his biographer Walter Isaacson, “Nobody is going to pay anything near that for something that looks like crap.” Musk believed that the way to start a car company was to build high-priced cars first and then let the technology trickle down to the mainstream. It was a classic Silicon Valley approach.
Alan Cocconi, the engineering whiz behind AC Propulsion, stands next to the company’s legendary tZero electric roadster in a picture taken in the early 2000s. The small yellow wheeled pod on the other side of the car is a trailer with a small gasoline engine that, when connected to the tZero, turned it into a hybrid gas-electric vehicle.Martin Eberhard
In Musk’s mind, it was all very obvious. He liked fast cars. He liked the tZero and believed it “could change the world.” He couldn’t even imagine why Gage was sitting here trying to sell him on the idea of the eBox. “Gage and Cocconi were sort of madcap inventors,” he told Isaacson. “Common sense was not their strong suit.”
Gage concluded that he wasn’t going to convince Musk to invest in AC Propulsion. “Well, if you want to do a sports car, then you should talk to Martin Eberhard,” Gage said. A few weeks later, Gage sent an email to Eberhard introducing him to Musk. “Elon Musk heads up SpaceX, is a car enthusiast,” Gage wrote. “He would be interested in hearing about your activities at Tesla Motors.”
As it happens, Eberhard and Marc Tarpenning, Eberhard’s partner and cofounder at Tesla, had considered contacting Musk even before Gage’s email arrived. They’d known of Musk and appreciated the way he thought. A few years earlier, they saw him speak at a Mars Society conference at Stanford University. Musk had talked about the rather improbable idea of sending mice to Mars. The presentation gave them a window into Musk’s unconventional approach to high-tech entrepreneurism and to life in general.
Eberhard and Musk agreed to meet, and then Eberhard emailed Gage. “Any chance of my borrowing the car for next week?” he wrote. Gage, of course, complied.
Martin Eberhard posed next to an electric motor at Tesla’s San Carlos, Calif., headquarters in 2006. Paul Sakuma/AP
By this time, Tesla was nine months old. It still had just three employees—Eberhard, Tarpenning, and Ian Wright, a New Zealand-born engineer and neighbor of Eberhard’s. The founders were arranging to pay the licensing fee on AC Propulsion’s drivetrain technology. And they were making arrangements to build their first cars using the chassis of the Lotus Elise two-seat roadster. They estimated they needed $6.5 million to go further. And that’s where Musk came in.
The original Tesla Roadster prototype, or “Mule,” was built inside the chassis of a 2002 Lotus Elise. Dylan Stewart/Image of Sport/Sipa/Alamy
Eberhard and Wright flew to Los Angeles on a Friday and met Musk in his cubicle at SpaceX. The meeting was supposed to last a half hour, but Musk’s questions came virtually nonstop, and as the meeting progressed, he repeatedly shouted to his assistant to cancel his next meeting.
Over the following weekend, Musk called Tarpenning to get his input about their financial model. “I just remember responding, responding, and responding,” Tarpenning said, according to a 2015 book by Ashlee Vance.
The Tesla founders were all impressed with Musk. He was unlike any of the VCs they’d met with in the previous months. He was technically astute. He’d earned a bachelor’s degree in physics from the University of Pennsylvania, and in his two-day stint as a Ph.D. student at Stanford, he’d intended to do a dissertation on solid-state capacitors for use in electric cars.
Moreover, he wasn’t averse to risk—at least not intelligent risk. He loved technical challenges, and he loved proving that the impossible was possible. “You’re presenting an electric car company to this person on the other side of the table, and he’s doing something even crazier,” Tarpenning said later. “He’s building rocket ships.”
On the Monday after their first meeting at SpaceX, Eberhard and Tarpenning flew back to Los Angeles. Musk agreed to invest $6.35 million. He would become the biggest shareholder as well as chairman of the company.
Now, Tesla Motors was really in business. All it needed was someone to design and build a groundbreaking electric car.
“All Electric Cars Have Sucked”
No one at AC Propulsion believed that Tesla Motors had even a remote chance of success. The whole idea—building and selling electric vehicles and competing against the giants in Detroit, Japan, and Germany—seemed impossible. Even Toyota, which was having so much success with the hybrid Prius, was not planning to build pure electric cars.
The prospect of starting any kind of auto company was unbelievably daunting. Automotive startups had been the undoing of many ambitious entrepreneurs, including Henry Kaiser, Preston Tucker, and John DeLorean. Such endeavors required mountains of money, connections, and expertise. There were unseen obstacles around every corner. And the people who’d launched Tesla, as smart as they were, were almost certainly unprepared for what lay ahead.
Years later, Musk would contend that their struggles were caused by the fact that Tesla had been founded on “two false premises.” The first was that Tesla’s founders believed they could simply convert an existing gasoline sports car to electric. The second was that they could use the existing AC Propulsion technology with little or no modification. “That turned out to be, in retrospect, staggeringly dumb,” Musk said.
In April or May of 2004, Tesla engineers worked on an early test vehicle, or “mule,” of the Roadster. The group included (clockwise from upper left) mechanical engineer Gene Berdechevsky, in the pink shirt, electrical engineer Phil Cole, and mechanical engineer Dave Lyons, in the dark blue shirt.Martin Eberhard
He also later concluded that their early path almost doomed them. “It ended up being much worse than if we had designed the car from scratch,” he said. But Tesla decided it could not go back and start over. It could only deal with the problem at hand.
Eberhard and Tarpenning were terrified of going into production with the existing analog motor controller and drivetrain electronics, which were unreliable and jittery. If those problems weren’t fixed, they knew their new vehicle would fail, and so would the company.
Another looming issue was the safety of the Tesla lithium-ion battery pack. To test it, Eberhard brought the engineering team to his home, where they dug a pit in his backyard. They took a brick of cells, covered it with a sheet of Plexiglass, and then remotely heated one of the cells with an electrical wire. As they expected, the heated cell burst into flame, setting neighboring cells on fire. The cells went off one at a time—pop, pop, pop. “We had a conflagration,” Eberhard said. “One cell caught fire, and it blasted right through the pack.”
Buyers of sports cars were known to be forgiving. In their quest for performance, they could put up with poor reliability. But the fire hazard was another matter, and one that had the potential to take down the company. Eberhard took the news right to Tesla’s board of directors. “It was my first big oh-shit moment to my board,” he said. “I told them we’ll have a day-to-day schedule stop until we figure this out.”
Working with friends from his Stanford days, Straubel began developing a new pack in his garage. The team acquired 7,000 lithium-ion cells from LG Chem, then constructed battery bricks, each with 69 cells, and tested them with different kinds of liquid-cooling systems. By October 2004, they’d finished a prototype pack and used a crane to lower it into the back of a Lotus Elise sportscar.
A few months later—in January 2005—the team had completed a working prototype of that first car. At the end of the month, they showed it off at a board meeting, and Musk took it for a spin. Impressed by its performance, he invested $9 million more, and Tesla completed a $13 million round of funding. Now the vehicle had a name—Roadster—and a tentative production schedule. The plan was to begin delivering it to customers in early 2006.
Tesla’s struggles with the Roadster were not apparent to the outside world, especially to those invited to the reveal of the Roadster at the Santa Monica Airport in July 2006. By then, the yellow test car, or “mule,” had evolved into two prototypes: a red car and a black car, both of which would be available for drives at the event.
Tesla unveiled the Roadster, its first vehicle, at the Santa Monica, Calif., airport on 19 July, 2006.Glenn Koenig/Los Angeles Times/Getty Images
The prototypes were more advanced than the mule, with more of a production-type design. But Musk and the team didn’t know what to expect at the reveal. The company was just coming out of stealth mode, and at that point, Tesla had received no media coverage. It had no customers, no deposits, and no sales team.
Still, Musk planned a huge party for the unveiling—an “awesome event,” in his words, staged inside the airport’s Barker Hangar. He told his personal assistant to invite 350 guests, including Michael Eisner of Disney, movie producer Richard Donner, actor Ed Begley Jr., California Governor Arnold Schwarzenegger, and many other luminaries. All were told to bring their checkbooks in case they wanted to write a $100,000 check to put a deposit on an electric Roadster. Meanwhile, a Roadster prototype zipped around a makeshift road inside the hangar, out the door, down a runway, and back inside again.
Musk took center stage, telling the audience that they were witnessing the start of a new era in automotive technology, according to CNET’s coverage of the event. “Until today, all electric cars have sucked,” he told the audience. “Electric cars play into the strength of Silicon Valley. A lot of the things inside the car are conventional automobile technology. The magic is the battery technology and the software and the controllers.”
Tesla Hooks Arnold Schwarzenegger, George Clooney
Musk’s message was perfect for such an event, especially for the dozens of reporters who were there to publicize the reemergence of the electric car. They adored the Roadster. It was small, powerful, electric, and above all, cool. It was anti-Detroit—a new kind of car that burned no gasoline and was born in Silicon Valley instead of an antiquated factory in Michigan.
The night was also a financial success for Tesla, with twenty $100,000 checks gathered from prospective buyers. And the momentum continued. A few days after the event, Joe Francis, creator of the adult entertainment franchise Girls Gone Wild, sent an armored truck to Tesla’s San Carlos office to drop off $100,000 in cash. A few days after that, Schwarzenegger put his money down, as did actor George Clooney. Within two weeks, Tesla had presold 127 Roadsters.
California Governor Arnold Schwarzenegger was among the first buyers of the Tesla Roadster on the day the car was officially unveiled, 19 July, 2006.Glenn Koenig/Los Angeles Times/Getty Images
Meanwhile, though, the company’s manufacturing woes continued. The mechanical problems weren’t even the biggest issue. The biggest issue was the supply chain. This was ironic, because some in the media admired Tesla for its global approach. They liked the fact that the battery pack, the motor, the chassis, and the assembly had an international flavor. It was a world car, they thought.
But for Tesla, it was a nightmare. The battery pack was being assembled in Thailand by a manufacturer of barbecue grills. The facility was 3 hours from Bangkok, literally in a jungle where the heat was almost unbearable, and the factory building consisted of a truss roof held up by some steel columns. There were no walls because no one there wanted to work indoors.
And because the pack assembler was inexperienced, Tesla engineers were repeatedly flying to and from Thailand to direct the effort. They would find animal droppings on the battery packs, which were sitting out in the open air all day and all night.
For the umpteenth time, Musk wondered if the company would be able to survive. “We’re doomed if we don’t in-source the battery pack,” he told one of Tesla’s manufacturing engineers, “because we have a supplier in Thailand who is great at making barbecues but not great at battery packs. And the supply chain is so long that it takes six months from when the cells are built to when the battery pack is done and in a car. So that means the capital cost is gigantic because we have to pay for all that inventory and process. And inevitably, there are mistakes in the design or fabrication of the battery pack, and then we have six months’ worth of battery packs that don’t work.”
Never mind that this chaotic approach was central to their plan. Tesla had never been envisioned as an old-fashioned, Detroit-style, vertically integrated manufacturer. From the beginning, it had been a Silicon Valley–type enterprise that would rely on others for the bulk of its manufacturing. Only now, as the fledgling company sent batteries and motors and assembled cars back and forth across two oceans, was its plan beginning to appear untenable. “We had this misguided idea that everything must be cheaper and better if built in Asia,” Straubel later said.
Tesla’s Chances of Success: 10 Percent
From the beginning, Musk had never been optimistic about Tesla’s chance of success. He repeatedly said he thought it was approximately 10 percent. “In 2004, the idea of starting a car company was extremely stupid,” he said. “The idea of starting an electric car company was stupid squared.” As he watched Tesla struggle with its supply chain, his earlier words were starting to look prescient.
The only chance, the engineering team concluded, was to bring the manufacturing of all of the subsystems, such as the battery packs, motors, and inverters, in-house. They disassembled their overseas operations and moved them to California, starting with battery pack manufacturing. Assembly stations were loaded into huge shipping containers, transported back to one of the company’s new facilities on Bing Street in San Carlos, and then reassembled there. It took five and a half months. They also redesigned the battery packs and developed machines for automating their assembly.
In 2008, with mass production of the Tesla Roadster just getting underway, Elon Musk gave an interview at the company’s headquarters, then in San Carlos in northern California. At the time, Tesla was merely a startup in a precarious position, bleeding cash and grappling with many manufacturing problems.Ryan Anson/Bloomberg/Getty Images
Musk concluded that the key to success was not the design of the car itself but rather the manufacturing. Henry Ford had reached the same conclusion a hundred years earlier. It was “the realization of how important it is to build the machine that builds the machine,” Musk said at the Tesla Annual Shareholder Meeting in 2016. “And how much harder it is to build the manufacturing system that builds the product, than it is to create the product in the first place. You can create a demo version of a product…with a small team in maybe three to six months. But to build the machine that builds the machine takes at least a hundred to a thousand times more resources and difficulty.”
Gradually, Tesla’s idea of letting others do its manufacturing slipped away. Packs were built in San Carlos, and then installed in the Lotus Elise chassis there instead of in England.
“We had control now,” said manufacturing engineer Jason Mendez. “We had all the engineers there. We didn’t have batteries on the water, not from Thailand to England and not from England to here.”
Musk began to talk about a new vision. He called it the gigafactory. Raw materials would enter at one end, and a car would exit at the other end. This was the ultimate in vertical integration, and it sounded a lot like Henry Ford’s vision for the River Rouge plant in Dearborn, Mich., in 1917.
Tesla Motors was becoming an auto manufacturer.
From Your Site Articles
Related Articles Around the Web
Tech
Inside the Inference Hardware Revolution Of 2026
Since about 2020, AI has largely focused on training bigger and better models. Large language models (LLMs) ballooned from millions of parameters to trillions. This proved effective: The largest version of OpenAI’s GPT-3, released in 2020, correctly answered just 43.9 percent of questions on a popular knowledge-and-reasoning benchmark. Just four years later, GPT-4o reached a score of 88.7 percent on the same exam, effectively matching those of human experts.
Advanced AI labs are still training ever larger models, but that training has somewhat receded to the background of the AI conversation. In 2026, inference—the use of trained models to produce code, write essays, or make images of ourselves as elves—has come to the forefront.
“It’s like training is yesterday’s news,” says Matt Kimball, principal data-center analyst at Moor Insights & Strategy. “All that any chief information officer wants to talk about is inference.” Nvidia CEO Jensen Huang, speaking at the company’s GTC 2026 conference, touted this change as the “inflection point of inference.”
Part of what’s caused the shift is very simple: LLMs are becoming useful, so people are using them. On top of that, many models on the market today are reasoning models. In response to a user’s query, they run inference not just once but multiple times, reprompting themselves in a process called chain of thought. Reasoning models generate longer outputs, and models with high reasoning effort can produce up to 20 times as much text as those with low or no effort. Adding even more to the world’s inference workload, the rise of agentic AI has resulted in inference running not just as a real-time response to a user’s query but also around the clock, working autonomously toward a user-defined goal.
Amazon’s Trainium chip was originally designed for AI training. However, Amazon Web Services chose to break up AI inference into two parts, with Trainium running the more computationally complex portion and Cerebras’s wafer-scale engine taking on the more memory-intensive portion.Amazon
The resulting explosion in inference demand has led to unexpected alliances among tech giants. OpenAI and Amazon have deployed chips the size of a dinner plate designed by Cerebras, despite Amazon having its own Trainium chips. Nvidia bought key talent and intellectual property from AI-inference startup Groq in a controversial deal worth US $20 billion. And Anthropic is paying LLM competitor SpaceXAI over a billion dollars per month to lease spare compute.
Although they might seem similar, AI training and AI inference are computationally different. These big moves from tech giants signal that in order to support the inference demand, we’re going to need a very different mix of hardware than experts may have expected even a couple of years ago.
How does AI inference differ from AI training?
An untrained LLM is like a jumble of Scrabble tiles on a table. Instead of single letters, though, the tiles show fragments of words, called tokens. Everything you’d need to write almost anything is present, but nothing makes sense.
Training a model organizes this jumble using a guessing game played at scale. The model is shown real text with the next token hidden and asked to predict what comes next. After each guess, the correct token is revealed and then compared to the prediction, and the difference is used to calculate the model’s accuracy. The game is played not with a single sentence but over billions of passages.
While a real game of Scrabble can be played over a bag of chips and a few drinks, AI training is computationally intense. The model updates its parameters through backpropagation, a process that repeatedly calculates how each of a model’s billions or trillions of parameters should shift to make the next prediction better. This is why tech giants are building larger data centers than ever before.
Eventually the model’s creator decides further training isn’t worth the cost, and the guessing game stops. Backpropagation ends, the parameters are frozen, and the LLM becomes a pretrained model. Fine-tuning—a short training run on smaller, more specialized data—adds final tweaks, and the model is deployed.
Nvidia’s Groq 3 language-processing unit minimizes data movement by placing on-chip SRAM memory and computational blocks in the order they are needed on-chip.
Nvidia
Next comes inference. This is the process of using the deployed model, which, now that it’s been trained, has learned to spit out Scrabble tiles—tokens—in a sensible order.
You might think that AI inference is less computationally demanding because the backpropagation calculations used to update parameters are eliminated. But Sudeep Bhoja, founder and CTO of the inference-hardware company d-Matrix, explains that inference adds new challenges.
The models are “autoregressive” in nature. That is, the next output depends on the previous one. “So to generate the next token, you have to read all of the weights and all of the [context] from the previous token,” explains Bhoja. The context includes all of your prompts, all of the LLM’s replies, and all of the files you upload. It’s a lot of data and a lot of processing.
An LLM generates its reply in two phases: prefill and decode. Prefill is the model reading a prompt. It processes every token at once, computing how each token relates to all the others. This operation is called attention, and it’s a defining characteristic of the transformer architecture behind modern LLMs. It allows them to respond to a word in its sentence, paragraph, and larger context rather than on its own. Think of it like arranging Scrabble tiles before you place them in a game. Many players move tiles around to imagine how they connect. Self-attention plays a similar role, though instead of moving physical tiles, each token sends a query to the others and receives a score indicating the token’s relevance.
These queries result in two types of vectors: the keys and values. They are typically placed in a store called the KV cache. This isn’t strictly required, as a model could instead recompute these vectors with each new token it generates. But nearly all LLMs use a KV cache to reduce how much computing they do. The KV cache is stored in memory and becomes a scratchpad to which the LLM can return to understand a conversation, and though it starts small, it can swell to dozens of gigabytes.
Prefill is a problem that can be easily divided up and worked on in parallel. This is why GPUs became the dominant AI accelerator as LLMs surged in popularity. Graphics rasterization (computing the color of every pixel on a screen) is also massively parallel, so GPU architectures were a natural fit.
Cerebras’s wafer-scale engine chips maximize memory bandwidth by keeping everything—both memory and computational units—side by side on the dinner-plate-size chips.
Cerebras
Next comes decode. Here, the model generates its reply one token at a time. At each step it takes the most recent token, weighs it against everything in the KV cache, uses that information to predict the next token, and adds the new token’s key and value to the cache. Then it repeats in sequence, token by token.
This is where the autoregressive nature of the model works against inference speed. Predicting each token requires reading the entire model from memory, and that model consists of possibly tens to hundreds of gigabytes of parameters (the numbers representing what the model learned in training). Crucially, this is in addition to the memory required to store the KV cache.
As a result, the movement of all this data through memory often requires more bandwidth than inference hardware has available. So at least some of the computing parts of a GPU sit idle as it waits for data. Researchers found that Nvidia H100 GPUs running open-source LLMs sit idle 50 to 80 percent of the time.
Memory’s role in inferencing
Shahriar “Sha” Rabii, former head of silicon engineering at Meta and cofounder of the AI startup Majestic Labs, says idled processors are why many companies that are trying to improve AI-inference performance are laser-focused on memory. “With the GPU-based approach, you end up greatly over-provisioning compute and starved on memory. That’s driving the big [memory] scale out,” he says.
Bhoja’s d-Matrix and Rabii’s Majestic Labs both focus on this memory bottleneck. However, their companies imagine different solutions.
d-Matrix’s second-generation AI accelerator, Raptor, aims to improve inference performance by minimizing the distance between compute and memory. The GPUs in most current AI-inference deployments do this by placing high-bandwidth memory (HBM) around the perimeter of the GPU. Each HBM is a stack of DRAM dies linked together and connected to a superfast interface to the GPU. This is great for training, but for inference, the amount of memory you can stack this way and the bandwidth it can provide leave something to be desired.
d-Matrix’s Raptor removes that bottleneck by stacking an AI accelerator on a DRAM die. Instead of stacking memory, d-Matrix stacks memory and compute. Bhoja says this reduces the distance that data must travel to “micrometers instead of millimeters.” Like building a skyscraper, going vertical makes it possible to do more inside the same physical footprint.
Majestic takes the opposite approach. Instead of trying to minimize the length that data must travel between compute and memory, the company is focused on improving the memory interface to accommodate longer wire traces while keeping bandwidth high. Longer wires allow Majestic to connect memory stacks that aren’t directly next to the GPU, removing the space limitation of HBM.
“A memory interface has a very short physical distance it can operate over. In the case of HBM, it’s up to 2 or 3 millimeters. You have this shoreline around the periphery, which is the only place where you can put HBM,” says Rabii.
Majestic claims its memory interface can transmit bits as far as about a meter. That’s achieved with a proprietary copper link and a memory-aggregator chip that coordinates data. “The aggregator is the endpoint for the high-speed interface and a way to fan out to many, many commodity DRAM chips,” says Rabii. Because of this, Majestic can support up to 128 terabytes of DRAM memory in a single server rack—a significant increase over Nvidia’s GB300 NVL72 rack, which has about 20 TB of HBM3E.
d-Matrix and Majestic have one thing in common: Instead of HBM, they both use off-the-shelf DRAM. This is the most common type of computer memory in the world; it’s in everything from smartphones to cars. Memory analyst Jim Handy says HBM costs two to three times as much as DRAM. d-Matrix and Majestic chose DRAM in part because of this price advantage. However, the proponents of HBM, which include memory giants like Samsung and SK Hynix, aren’t sitting idle.
HBM4, the latest version of HBM memory, is now in production and will be used by Nvidia’s Vera Rubin GPU, which is expected to ship in the second half of 2026. Hoshik Kim, head of memory-systems research at SK Hynix, says HBM4 “will decisively break the memory bottlenecks constraining AI inference today” by doubling HBM’s maximum memory bandwidth and increasing the amount of HBM memory per stack.
Combining chips for faster inference
The big players—Nvidia and Amazon—are going for an all-chips-on-deck approach. Nvidia’s GPUs and Amazon’s Trainium training accelerators are still great for part of the inference workload: the prefill stage, where all the context keys and values are calculated. But to accelerate decode, the part where new tokens are generated, they are looking to new, memory-centric architectures from smaller players.
In Nvidia’s case, the smaller player was Groq (not to be confused with Grok, the family of LLMs trained by SpaceXAI). Nvidia purchased intellectual property and hired talent from Groq at the end of 2025, and just three months later at the Nvidia’s GTC 2026 conference, Jensen Huang unveiled the Nvidia Groq 3 language-processing unit (LPU). Groq’s architecture relies on memory—in its case, SRAM—built directly into the chip’s architecture.
Unless you’re a chip architect, or a hardcore PC gamer, you probably never give SRAM a thought. SRAM has the benefit of being tightly integrated into a compute chip’s architecture—it’s on the same piece of silicon as the processor—and has the drawback of being less dense and more expensive than DRAM. Most chips include only a few dozen megabytes of SRAM. AI inference, however, has ignited new interest in SRAM as a means of bringing the model weights stored in memory closer to compute.
Ian Buck, vice-president and general manager of hyperscale and high-performance computing at Nvidia, says the LPU has a much different set of priorities than the company’s GPUs. The LPU has far less raw computing power than a standard GPU, but it gains 500 megabytes of on-die SRAM connected directly to its floating-point math units. “The benefit is the memory bandwidth. The LPU has seven times the memory bandwidth of the GPU,” he says.
Between the Rubin GPU and the Groq LPU, prefill and decode can both be accelerated to get the best of both worlds, the theory goes. “We do all the attention math and context processing on the Vera Rubin [GPU] rack,” explains Buck. “For all the expert calculations…the matrix multiplications, we do that part on the LPU.” The company packs 256 LPUs into the Groq 3 LPX, a system the size of a data-center rack.
Amazon Web Services (AWS), for its part, struck a deal with Cerebras, to pair the Trainium accelerator with Cerebras’s Wafer-Scale Engine 3 (WSE-3). Cerebras takes a similar approach to Groq, though at a much larger scale. WSE-3 turns an entire silicon wafer into a single chip that contains over 4 trillion transistors. The design doesn’t connect to external memory but instead etches 44 gigabytes of SRAM into each wafer. “We store the [model] weights on the SRAM,” says James Wang, formerly director of product marketing at Cerebras who has since moved to SpaceXAI. “So that’s easily 40 to up to 80 billion parameters that we can support on one chip.”
Amazon plans to use AWS Trainium chips for prefill, and Cerebras for decode. But Cerebras’s chips can also go it alone in inference. WSE-3 was deployed by OpenAI to power GPT-5.3-Codex-Spark, a variant of the company’s coding mode, outputting over 1,000 tokens per second. For comparison, OpenAI’s standard GPT-5.4 deployment outputs 50 to 125 tokens per second.
Cerebras can also tackle prefill without moving the workload to different specialized chips. For this, it networks together multiple WSE-3 chips to form a single pool of memory. “Commercially, we’ve done about 500 billion parameters for our customers up to this point,” says Wang. “But the architecture has no innate limitation in terms of how many parameters it will do.”
Despite these differences in strategy, Nvidia and AWS seem to agree that the future of AI inference will be solved by a systems approach that pools different kinds of chips together to tackle the largest LLMs. Or, as Buck says: “To do modern AI inference, you need all the chips.”
Learning to do more with less (bits)
Nvidia became the world’s most valuable tech company because it designed the world’s most desired GPUs. But not all of the attention is focused on improving AI-inference hardware. AI researchers are also learning how to optimize LLM software and hardware in tandem to make the best use of the memory and compute components.
Most computers store numbers in a 32-bit or 64-bit format. These determine how many bits are available to represent a single number. If too few bits are available, the number can’t be stored without losing information. The quality of an LLM benefits from more-precise number formats, but this creates a problem for inference performance. More-precise numbers aren’t free. The bits that describe them take up more space in memory and require more silicon and energy to compute.
Gilles Backhus, cofounder of the AI-accelerator company Tensordyne, says this creates a tension between model size and number precision. “Would you prefer a model that is size x but runs in 8-bit, or would you prefer a model that is twice the size but runs in 4-bit?” The size of each model will be roughly the same in terms of memory and compute, “but the 4-bit approach gives you twice as many synapses, if you will. And people are figuring out that [the 4-bit approach] is worth it.”
The process of converting an LLM from a more-precise number format to a less-precise format is called quantization, and it’s been in use for several years. However, researchers are finding new ways to quantize models down while retaining a large majority of the model’s quality.
Nvidia recently created a new 4-bit number format, NVFP4, for this purpose. AMD, Intel, and Qualcomm have instead rallied around a competing 4-bit number format called MXFP4 that Nvidia also contributed to developing. “It’s the black art of AI,” says Buck, of Nvidia. When Nvidia quantized DeepSeek-R1 from FP8 to NVFP4, scores on seven major benchmarks degraded by less than one percent while performance improved by three times, the company says.
Quantization is likely just the tip of the spear, as AI researchers and startups are investigating a diversity of opportunities for optimization, some of which could dramatically change the silicon found in AI-inference hardware.
Tensordyne’s unique approach to AI inference combines a logarithmic number format with bespoke hardware in the company’s Napier chip. Tensordyne
Tensordyne is expected to accelerate AI inference with a logarithmic number system that leans on a property of logarithms: The log of A times B equals the log of A plus the log of B. So, storing numbers as their exponents lets the chip add where it would otherwise multiply. That matters in silicon because multiplier circuits draw more power and use more die area than adders do. Tensordyne says its rack-scale hardware, called Napier, can produce up to 1,300 tokens per second per user, and can do so while using less than a tenth as much power as comparable Nvidia hardware.
Etched, a startup based in San Jose, Calif., is even designing AI accelerators that translate the transformer architecture used by LLMs directly into silicon. Rather than building general-purpose GPUs, the company is wiring up the connections needed for efficient transformer calculations into its chip, making the chip much less flexible but more efficient for the tasks most performed by current LLMs. The company says its first AI accelerator, Sohu, can run Meta’s Llama 70B model at a stunning 500,000 tokens per second, though this approach also means it won’t be able to run LLMs that move away from a typical transformer architecture.
Whether these ideas will prove fruitful remains to be seen. Etched just shipped their first rack in August. Tensordyne believes its first hardware will be available in 2027. Even so, these startups show how the demand for inference performance is fueling unconventional ideas.
Inference is everyone’s game
The sheer variety of approaches to AI-inference acceleration—stacking compute on memory, extending interfaces from millimeters to meters, using an entire silicon wafer for SRAM, squeezing models into 4 bits—raises a question: Which is going to win, and which is going to lose?
But that’s likely not the right question, experts say. The demand for AI is currently insatiable, and while fears of an AI bubble stalk the industry, it has yet to hamper growth.
On the contrary, Kimball of Moor Insights & Strategy thinks inference could drive intense demand for AI hardware in the long term, because it’s not obvious where that demand will end. “You could add a million agents into your organization,” he says. “These things work 24 hours a day; they don’t go home at five at night like we do.”
If AI inference remains as desirable as Kimball expects, the evolution is likely to follow the same trajectory as the CPU. The CPU didn’t improve along a single axis but instead across multiple fronts simultaneously. Once transistor scaling slowed, chip and system architecture innovations of all kinds proliferated. The list of individual innovations that led to today’s ubiquitous, powerful personal compute could fill dozens of books.
A few decades from now, the history of AI inference innovation will show similar depth.
From Your Site Articles
Related Articles Around the Web
-
Fashion4 days agoWeekend Open Thread – Corporette.com
-
Business6 days agoMicron Stock Climbs Above $1,031 as AI Memory Crunch and a $50 Billion Outlook Fuel the Rally
-
Tech2 days agoThe Latest Weird Thing to Play Doom Is the Mapped-Out Brain of a Fruit Fly
-
Business6 days agoAMD Stock Climbs After Management Lifts 2027 Data Center Outlook Toward $70 Billion in AI Sales
-
Crypto World7 days agoBitcoin price risks $76K drop as $78K support weakens
-
Crypto World4 days agoXAG/USD: Silver’s Short-Term Rally Meets Its Moment of Truth
-
Crypto World5 days ago2 Chip Stocks Broke Out This Week. Neither Was Nvidia
-
Business4 days ago10 Most-Streamed Songs On Spotify In 2026 So Far, Led By Ella Langley’s Dominant Run On The Charts This Year
-
Crypto World7 days agoEthereum price stalls below $2,500 as ADX drops to 11
-
Tech5 days agoBattery life is the only iPhone 18 Pro and iPhone Duo upgrade I care about. Apple didn’t disappoint
-
Crypto World5 days agoOKX launches 10x OpenAI, Anthropic X-Perps in Europe
-
Crypto World6 days agoPi Network ships Protocol 27 on a network with 14 million users and zero DeFi
-
Crypto World4 days agoDiesel Tops $6 a Gallon for the First Time as 28 States Set Records
-
Crypto World7 days agoCLARITY Act runs out of calendar as crypto regulation stalls
-
Tech6 days agoApple Watch Ultra 4 vs Watch Ultra 3: Should you really spend another $799?
-
News Videos4 days agoFacing Financial Fears
-
Crypto World6 days agoBitcoin price risks $70K if $78K neckline breaks
-
Crypto World6 days agoBitcoin price holds near $79K as cycle drawdowns narrow
-
NewsBeat7 days agoWhat went right this week: an ‘historic’ fall in violent crime, plus more
-
Business7 days agoMeta debuts long-awaited personal AI agent, Muse



You must be logged in to post a comment Login