Connect with us

Tech

BYD will unveil its first humanoid robot in August

Published

on

BYD sells more electric cars than anyone. Next month, it wants to sell you one using a robot.

China’s biggest EV maker will unveil its first humanoid robot in August, at its Di Space experience centres in Zhengzhou, the company told the South China Morning Post. It will be a working prototype, not a concept, and it will mingle with visitors.

The robot reportedly has a name and a job. According to Chinese outlet KrASIA, citing a since-deleted BYD post, it is called “Xiao Di,” a service humanoid that stands 1.61 metres, weighs 58.5kg and can translate between six Chinese dialects and six foreign languages in real time. BYD has not officially confirmed the specs.

Cars first, then everything else

The plan starts in the showroom. Executive vice-president Stella Li wants two or three robots in every BYD store, greeting customers and explaining cars. She insists they will assist human staff, not replace them.

Advertisement

The 💜 of EU tech

The latest rumblings from the EU tech scene, a story from our wise ol’ founder Boris, and some questionable AI art. It’s free, every week, in your inbox. Sign up now!

From there, the ambition widens: supermarkets, malls and warehouses in the medium term, and eventually homes, doing the cleaning, cooking and companionship. BYD says it will build an open platform that makes both its own robots and models co-developed with others.

Its pitch is manufacturing. BYD already builds its own batteries, motors and electronics at huge scale, and Li argues cars and robots share the same roots. That, in theory, lets it build robots more cheaply than a standalone startup can.

Advertisement

Everyone with a car factory wants one

BYD is not early. It is late. Tesla’s Optimus is the headline rival, but China’s carmakers are swarming in. Xpeng is trialling its Iron robot, Li Auto is exploring designs, and Chery’s Aimoga is already selling to consumers.

The backdrop is a humanoid boom. China made about 20,000 humanoids in 2025 and more than 40,000 in the first half of 2026 alone, with the government pushing for far more. Some forecasts stretch to tens of millions of robot workers within a decade.

The catch, and the timing

There is a reason for the robot rush, and it is not entirely rosy. BYD’s car business is under strain, with first-half sales down about 16% amid China’s brutal price war. A new growth story is welcome.

But a slick demo is not a business. Investors still have no price, no production timeline and no paid deployments to judge. BYD has also denied separate reports of a factory robot line, so much remains unconfirmed.

Advertisement

And the overseas door is closing. The move lands just as the US moves to bar imports of Chinese robots, which could keep Xiao Di out of one of its biggest potential markets. For now, BYD’s robot has one job: sell cars, in China, from a showroom floor.

Source link

Advertisement
Continue Reading
Click to comment

You must be logged in to post a comment Login

Leave a Reply

Tech

AI price wars: OpenAI cuts GPT-5.6 Luna prices by 80% as model competition shifts toward cost

Published

on

To quote an ancient Jedi Master “Begun, the AI price wars have!”

OpenAI is sharply reducing the prices of two models in its GPT-5.6 frontier series, cutting GPT-5.6 Luna, the smallest and fastest model in the series, by 80% and GPT-5.6 Terra, the mid-tier model, by 20%, while adding a premium Fast mode for its flagship GPT-5.6 Sol model.

The cuts place Luna much closer to the lowest-cost commercial models in the market and arrive just a few days after Anthropic released its highly performant Claude Opus 5 at the same price as Opus 4.8, and Google introduced Gemini 3.6 Flash and Gemini 3.5 Flash-Lite, two rival models built around lower inference costs, faster execution and more efficient agent workloads.

OpenAI is successfully undercutting Google’s price per intelligence and attempting to sway Anthropic users, who may not mind paying more, with a speed boost.

Advertisement

OpenAI says Luna will now cost $0.20 per million input tokens and $1.20 per million output tokens, for a combined input-plus-output price of $1.40 per million tokens.

Terra will cost $2 per million input tokens and $12 per million output tokens, for a combined price of $14.

Pricing for Sol Standard remains unchanged at $5 per million input tokens and $30 per million output tokens. OpenAI is also adding Sol Fast mode at twice the Standard price: $10 per million input tokens and $60 per million output tokens.

The company says Fast mode delivers up to 2.5 times the throughput without changing the model’s underlying intelligence.

Advertisement

OpenAI co-founder and CEO Sam Altman took to X to announce the changes as “major price cuts today.”

VentureBeat Frontier AI model API pricing comparison

Model

Input ($/1M)

Output ($/1M)

Advertisement

Total ($/1M)

Source

MiMo-V2.5 Flash

$0.10

Advertisement

$0.30

$0.40

Xiaomi

deepseek-v4-flash

Advertisement

$0.14

$0.28

$0.42

DeepSeek

Advertisement

deepseek-v4-pro

$0.435

$0.87

$1.305

Advertisement

DeepSeek

GPT-5.6 Luna

$0.20

$1.20

Advertisement

$1.40

OpenAI

MiniMax-M3

$0.30

Advertisement

$1.20

$1.50

MiniMax

LongCat-2.0 — limited-time promo

Advertisement

$0.30

$1.20

$1.50

LongCat

Advertisement

Gemini 3.1 Flash-Lite

$0.25

$1.50

$1.75

Advertisement

Google

Qwen3.7-Plus

$0.40

$1.60

Advertisement

$2.00

Alibaba Cloud

MiMo-V2.5

$0.40

Advertisement

$2.00

$2.40

Xiaomi

Gemini 3.5 Flash-Lite

Advertisement

$0.30

$2.50

$2.80

Google

Advertisement

LongCat-2.0 — standard

$0.75

$2.95

$3.70

Advertisement

LongCat

MiMo-V2.5 Pro (≤256K)

$1.00

$3.00

Advertisement

$4.00

Xiaomi

GLM-5.2

$1.40

Advertisement

$4.40

$5.80

Z.ai

Grok 4.5

Advertisement

$2.00

$6.00

$8.00

xAI

Advertisement

MiMo-V2.5 Pro (>256K)

$2.00

$6.00

$8.00

Advertisement

Xiaomi

Gemini 3.6 Flash

$1.50

$7.50

Advertisement

$9.00

Google

Qwen3.7-Max

$2.50

Advertisement

$7.50

$10.00

Alibaba Cloud

Gemini 3.5 Flash

Advertisement

$1.50

$9.00

$10.50

Google

Advertisement

Gemini 3.1 Pro Preview (≤200K)

$2.00

$12.00

$14.00

Advertisement

Google

GPT-5.6 Terra

$2.00

$12.00

Advertisement

$14.00

OpenAI

GPT-5.4

$2.50

Advertisement

$15.00

$17.50

OpenAI

Kimi K3

Advertisement

$3.00

$15.00

$18.00

Moonshot AI

Advertisement

Gemini 3.1 Pro Preview (>200K)

$4.00

$18.00

$22.00

Advertisement

Google

Claude Opus 5

$5.00

$25.00

Advertisement

$30.00

Anthropic

GPT-5.5

$5.00

Advertisement

$30.00

$35.00

OpenAI

GPT-5.5 Instant (chat-latest)

Advertisement

$5.00

$30.00

$35.00

OpenAI

Advertisement

Sakana Fugu Ultra (≤272K)

$5.00

$30.00

$35.00

Advertisement

Sakana AI

GPT-5.6 Sol — Standard mode

$5.00

$30.00

Advertisement

$35.00

OpenAI

Claude Fable 5 / Claude Mythos 5

$10.00

Advertisement

$50.00

$60.00

Anthropic

GPT-5.6 Sol — Fast mode

Advertisement

$10.00

$60.00

$70.00

OpenAI

Advertisement

Pricing is shown per one million tokens. Total cost is calculated as input price plus output price. Cached-input pricing is excluded to keep the comparison consistent across providers.

OpenAI moves Luna into the low-cost tier

The most consequential change is the Luna price cut.

When OpenAI introduced the GPT-5.6 series, Luna was priced at $1 per million input tokens and $6 per million output tokens, for a combined total of $7. The new pricing reduces that combined figure to $1.40.

That places Luna below Google’s Gemini 3.5 Flash-Lite, which costs a combined $2.80 per million input and output tokens, and far below Gemini 3.6 Flash at $9. Luna also now costs less than OpenAI’s own GPT-5.4 and Terra models by a wide margin.

Advertisement

It is not the cheapest model in the broader market. Xiaomi’s MiMo-V2.5 Flash, DeepSeek’s flash model and several other APIs remain less expensive on a pure token basis. But the reduction brings an OpenAI frontier-series model into direct competition with the market’s low-cost inference tier.

OpenAI says the GPT-5.6 series represents its frontier model family, with Sol positioned at the top of the lineup, Terra as the middle tier and Luna as the smallest and fastest option.

The lineup was initially released in late June 2026 through a limited rollout by U.S. government request, before broader access, with each model intended to offer a different tradeoff among intelligence, latency and cost.

Sol is aimed at the most complex reasoning-heavy and agentic workloads, including advanced coding, multi-step planning and tool-using systems, while Terra is designed for general production use where a balance of capability and efficiency is required. Luna is positioned for high-throughput, low-latency tasks such as summarization, classification, routing, and lightweight real-time assistants where cost per request is the primary constraint.

Advertisement

Terra drops to match Google’s Gemini 3.1 Pro pricing

Terra’s 20% reduction moves its combined price from $17.50 to $14 per million tokens.

At that level, Terra now matches Google’s Gemini 3.1 Pro Preview pricing for context windows of 200,000 tokens or less.

It also undercuts OpenAI’s GPT-5.4, which remains priced at $2.50 per million input tokens and $15 per million output tokens, offering the same intelligence for about 1/13th the cost, as Krea AI’s Nic Dunz noted on X:

The adjustment creates a wider separation between OpenAI’s three GPT-5.6 tiers. Luna costs one-tenth as much as Terra on a simple combined input-plus-output basis, while Terra costs 60% less than Sol Standard.

Advertisement

Sol Fast moves in the opposite direction. At a combined $70 per million tokens, it is the most expensive model configuration in the comparison below, reflecting OpenAI’s decision to charge a premium for latency-sensitive workloads rather than lower Sol’s base price.

Cuts follow Google’s low-cost Gemini releases and Anthropic’s Claude Opus 5

OpenAI’s pricing changes come only about a week and a half after Google introduced its own low-cost Gemini 3.6 Flash and Gemini 3.5 Flash-Lite.

Google priced Gemini 3.6 Flash at $1.50 per million input tokens and $7.50 per million output tokens. Gemini 3.5 Flash-Lite costs $0.30 per million input tokens and $2.50 per million output tokens.

Google framed both models around the economics of agent deployment, arguing that lower token usage, fewer reasoning steps and reduced tool calls could lower the total cost of long-running software engineering and knowledge-work tasks.

Advertisement

Gemini 3.6 Flash reportedly uses 17% fewer output tokens than Gemini 3.5 Flash on the Artificial Analysis Index, with savings reaching as high as 65% on some long-horizon engineering workloads. Gemini 3.5 Flash-Lite is positioned as the fastest model in Google’s 3.5 series.

However, OpenAI’s models are more performant than Google’s, according to third party analysis outfits like Artificial Analysis, with even the Luna model outperforming Gemini 3.6 Flash and the older Gemini 3.1 Pro model, making the cost-per intelligence much more favorable to OpenAI.

Artificial Analysis Intelligence Index July 2026 snapshot

Screenshot of Artificial Analysis Intelligence Index as of July 2026

As AI coding startup Cognition noted on X, GPT-5.6 now “sits on the pareto curve of price/performance efficiency,” posting an animation of the GPT-5.6 series moving left on a chart representing intelligence on the y axis and cost on the x, showing that the models now offer among the most superior intelligence for lowest cost on the market.

Advertisement

And yet, rival Anthropic’s Claude Opus 5 remains about as performant as GPT-5.6 Sol, yet is 6% cheaper.

The model costs $5 per million input tokens and $25 per million output tokens—the same rates as Opus 4.8—but Anthropic says it delivers nearly all the intelligence of its more expensive Fable 5 model at roughly half the cost.

Unlike OpenAI’s Luna and Terra changes, Anthropic did not reduce the Opus API sticker price. Instead, it effectively lowered the price per unit of capability by replacing Opus 4.8 with a more capable model at the same $30 combined input-and-output rate. Anthropic also added an adjustable effort setting that allows developers to trade reasoning depth for speed and token savings.

That distinction matters for enterprise buyers. OpenAI is directly cutting per-token rates, Google is pairing lower prices with reductions in token use and tool calls, and Anthropic is emphasizing stronger task performance at an unchanged price. All three approaches target the same operational metric: the total cost of completing production work, rather than the advertised cost of an individual token alone.

Advertisement

The timing highlights how quickly pricing has become a competitive lever among frontier model providers. OpenAI’s response does not introduce a new model generation. Instead, it changes the economics of deploying models that were released only recently.

The market shifts from model access to model economics

The cuts indicate that access to frontier-level capability is no longer the only point of competition. The next question for enterprises is how cheaply and predictably those models can run in production.

OpenAI is still not the lowest-priced provider on a pure token basis. But Luna’s 80% reduction materially changes its position, moving it from the middle of the market into a pricing tier populated by smaller models from Google, Xiaomi, DeepSeek, MiniMax and other vendors.

That matters most for high-volume applications, where relatively small differences in token pricing can compound across coding agents, document systems, internal search tools and automated workflows.

Advertisement

OpenAI’s latest move therefore looks less like a routine adjustment and more like a repositioning of the GPT-5.6 series. Sol remains the premium option, Terra moves closer to competing pro-tier systems, and Luna becomes the company’s direct answer to the industry’s growing low-cost model segment.

Source link

Continue Reading

Tech

Banning Open-Source AI Models to Protect Our Cybersecurity May Do the Opposite

Published

on

In an escalating effort to give the federal government power over the AI industry, some members of the Trump administration have reportedly tried to implement a “de facto ban” on foreign-made open-source AI models. This ban would apply to any non-US-made AI model, but the goal seems to be to specifically target Chinese AI labs, which often release their models as open source. 

The recent release of the Kimi K3 AI model from Chinese developer Moonshot fueled these concerns. Kimi K3 matched and in some cases exceeded the capabilities of American-made AI models such as those made by OpenAI, Anthropic and Google. Its July release sent shockwaves through Wall Street – not unlike other AI model drops, but with one big difference. Kimi K3 was released as an open-weight model, while American AI leaders like OpenAI and Anthropic companies keep their tech closed with very few exceptions.

Open-source AI typically refers to open-weight models. Weights are characteristics that tell the model how to behave – giving more weight to useful answers than incorrect ones, for example. Open-source advocates have said that to be truly open-source, AI companies should release or clarify their training data sources. AI companies haven’t been forthcoming; OpenAI and other companies are being sued by publishers and artists for allegedly violating their copyrights in AI training. (Disclosure: Ziff Davis, CNET’s parent company, in 2025 filed a lawsuit against OpenAI, alleging it infringed Ziff Davis copyrights in training and operating its AI systems.)

Open-weight models give developers more insight into how top models work. Nearly 80% of developers use open models, a recent Mozilla report found. “Open-weight models are everywhere in the industry already,” said Linda Griffin, vice president of global policy at Mozilla. “So a world without them would hit a lot of people.”

Advertisement

Tech companies immediately and strongly rejected the idea of a government ban on open AI models. Microsoft, Nvidia, Meta and several other AI developers and tech venture capital firms signed an open letter (PDF) asking the Trump administration to reconsider. 

“Our AI leadership will be judged not by one frontier AI model, but by whether the United States builds a strong, open ecosystem that diffuses into every sector,” the July 24 letter reads. Ideally, the money and power that AI companies promise come with AI adoption, would follow.

Restrictions on Chinese tech: Good or bad for cybersecurity?

This is far from the first time the Trump administration has taken steps to limit Americans’ access to Chinese tech. Remember the years-long battle over potentially banning TikTok? Chinese parent company ByteDance was eventually forced to transfer ownership of its US business to US-based ownership led by Oracle, whose co-founder and executive chairman, Larry Ellison, is a prominent supporter of President Donald Trump.

Restrictions on foreign hardware have been rolling out, too. The Federal Communications Commission banned foreign-made routers in March, saying that they posed a security risk. The order massively disrupted the industry behind the devices that people need in order to access the internet.

Advertisement

Given the leaps in AI advancement over the past year, it isn’t totally surprising to see these arguments. AI is being used by both cybersecurity attackers and defenders, making it a high priority for AI companies to secure their models against potential misuse. 

Anthropic and OpenAI have both worked with the US government to slow-roll the release of their newest models, Claude Fable 5 and GPT-5.6, respectively. That government review is voluntary for now, but it might one day become mandatory, specifically because of the cybersecurity concerns.

But if government officials are worried about AI cybersecurity, banning open AI models might have the opposite effect.

“Open-weight models allow researchers to examine how these systems work and identify risks and vulnerabilities,” said Aditya Vashistha, professor of computer science at Cornell University. “Restricting access would make it much harder to independently evaluate how safe and secure these technologies really are.”

Advertisement

The Trump administration has been adamant that it won’t hinder AI innovation with regulation. But if open-weights models are banned, the people and companies who aren’t part of selective AI cybersecurity programs like Anthropic’s Project Glasswing will be at risk, said Ayham Boucher, executive director of AI strategy and innovation at Cornell.

“Regulators can’t have it both ways. You can’t restrict access to frontier models and ban open-source models, leaving enterprises and institutions defenseless,” Boucher said.

A logo against a black background with the words Project Glasswing next to a square with an abstract webbed pattern in it.

The AI company Anthropic created a consortium of tech companies that includes Apple, Nvidia and Amazon AWS to address the issue of cybersecurity in the era of next-generation artificial intelligence models.

Anthropic

Openly ‘diffusing’ the benefits of AI

Unsurprisingly, the tech industry has reacted negatively to the idea of restricting access to open-source AI models. Meta CEO Mark Zuckerberg advocated for open AI technologies, writing in an op-ed in the Wall Street Journal earlier this week that it is through openness that the benefits of AI spread throughout our society.

“Historically, hoping that an absolute power will benevolently provide for humanity if sufficiently enlightened hasn’t led to safe or positive outcomes,” Zuckerberg wrote. (As CEO of the company that operates several of the world’s largest social media platforms, Zuckerberg himself is one of the very few who have something like absolute power over our technological experiences.)

Advertisement

But the true appeal and benefit of open-source AI, like all open-source technologies, is that anybody gets to build with AI, not just Big Tech. 

“If you take them away, developers and consumers pay more, get less choice, and a whole lot of useful stuff just never gets made,” Griffin said. A broad ban “would set a troubling precedent.”

Developers know that some guidance is necessary. Over 1,000 staffers at top labs, including chief scientists from OpenAI, Meta and more, signed an open letter asking the US government to support an international effort to build technical and governance tools as they “pace the frontier” of AI research and development.

“For the US to maintain its leadership in AI, it cannot rely solely on closed models,” said Vashistha. “If the US moves away from open models while others continue investing in them, it risks ceding not just market share, but also influence over the global AI ecosystem.”

Advertisement

Source link

Advertisement
Continue Reading

Tech

LinkedIn realizes its users have been bathing in AI slop, offers a shower

Published

on

ai and ml

AI-riddled social network adds button to report sloppy posts, ditches AI rewrite tools, and promises more to come

LinkedIn has been drenching its users in AI-powered slop posts, going so far as to encourage writers to trade their own voices for a bot’s by hitting the “enhance your post” button when they want to share thoughts.

Enough is enough. On Thursday, the site’s Chief Product Officer has announced several changes designed to rehab the platform’s reputation. 

Advertisement

CPO Hari Srinivasan took to the Microsoft-owned social network following reports Thursday that LinkedIn had introduced a button for users to flag posts as AI slop. He confirmed not only that the reports are real, but that LinkedIn had additional plans as well.

The “Seems like AI slop” button is being added to the ellipsis menu on LinkedIn posts now, a spokesperson told us, and will be available on all posts and comments. Srinivasan said this button will not only allow users to report posts with sentences written like this – it will also help LinkedIn refine its AI-spotting AI models to help improve user feeds.

In addition, the “enhance your post” button, which used AI to tweak your wording, is being replaced by an option to have AI proofread your work while leaving your voice intact, instead of stealing it like the sea witch Ursula.

There’s been no shortage of scorn from The Register and elsewhere about LinkedIn’s rapid decline into a slop tank filled with faux thought leadership posts and generative drivel. As far back as 2024, reports were coming out that more than half of long-form LinkedIn posts longer than 100 words were believed to be AI generated. That hasn’t changed in the past two years.

Advertisement

“AI slop is a top priority for all of us. We really care about this,” Srinivasan wrote. “People come to LinkedIn to connect with real people and share their real perspectives,” he said, adding that LinkedIn wants to keep it that way – or, more realistically, nudge the platform back toward its former state.

Coincidentally or not, Originality.ai, the same AI detection site behind the 2024 report, released an updated scan of LinkedIn posts longer than 100 words on Thursday. According to this new report, a full 81 percent of the 5,000 posts it looked at this month were classified as likely being AI-generated. 

Srinivasan said it’s not just users employing AI to generate slop – it’s AI automating garbage posts and comments at scale, throughout the site.

“Everyday we are now catching hundreds of thousands of automated comment attempts, and have blocked billions of other automation attempts (posting at scale, slop) in the last couple months alone,” the LinkedIn CPO said. 

Advertisement

To that end, the professional social network is also “ramping up a series of new and improved classifiers that identify if a post is AI-slop or generally low-quality content,” Srinivasan said.

User analytics dashboards will also be getting a new feature that will tell posters when members flag their posts as potentially being AI, as Srinivasan said LinkedIn wants users “to get feedback from real humans on what sounds authentic – not just have an AI detector review it and get it wrong.”

“AI and slop are not the same thing,” Srinivasan added. “Many people refine thoughts with AI, and we believe they want to know when they sound inauthentic.” ®

Source link

Advertisement
Continue Reading

Tech

Thinking Machines debuts Inkling Small open source AI model nearing performance of predecessor at about 1/4 size

Published

on

Just two weeks after Thinking Machines released Inkling, its first open source AI language model, the well-funded startup led by former OpenAI chief technology officer Mira Murati today introduced Inkling-Small without sacrificing much of any performance — and in fact, the new model surpasses its larger predecessor on several benchmarks.

Inkling Small is a 276-billion-parameter multimodal reasoning model with a permissive Apache 2.0 license that comes within a single point of its larger sibling on the third-party Artificial Analysis Intelligence Index, despite the original Inkling being 975 billion parameters (internal model settings). It accepts text, image and audio inputs, produces text, and supports a context window of up to one million tokens.

Inkling Small uses 12 billion active parameters per token, compared with Inkling’s 41 billion active parameters, while preserving much of the flagship’s coding, reasoning and multimodal performance.

For enterprises, the appeal is not simply that Inkling-Small is smaller. It is that developers appear to give up relatively little capability while reducing the model’s compute requirements, inference costs and deployment footprint.

Advertisement

The model remains far too large for a laptop or conventional workstation, but it is materially easier to operate than the 3.5X larger flagship, making it a good fit for enterprises with some — but not a lot — of their own graphics processing units (GPUs).

Thinking Machines has released the full weights on Hugging Face and added support for fine-tuning through its Tinker model training application programming interface (API).

At launch, the company is advertising a limited-time 50% discount, bringing API pricing for the standard 64K-context Inkling-Small model to $0.58 per million prefill (input) tokens, $1.44 per million sampled (output) tokens, and $1.73 per million training tokens, with cached prefill requests priced at $0.116 per million tokens. A 256K-context variant is also available at higher rates.

Nearly the same performance at a quarter the size

Artificial Analysis assigned Inkling-Small a score of 40 on its Intelligence Index, compared with 41 for Inkling.

Advertisement

That result is notable because Inkling-Small has 276 billion total parameters and 12 billion active parameters, while Inkling has 975 billion total parameters and 41 billion active parameters.

Artificial Analysis also reported that no open-weight model at Inkling-Small’s size or smaller scored higher on the index.

The model does more than merely approach the flagship’s aggregate score. On several evaluations, it surpasses Inkling.

Thinking Machines reports that Inkling-Small scores 80.2% on SWE-bench Verified, compared with Inkling’s 77.6%, and 64.7% on Terminal Bench 2.1, compared with 63.8% for the larger model. It also edges ahead on SciCode, Humanity’s Last Exam, GPQA Diamond and CritPt.

Advertisement

The gains are not universal. Inkling retains a clear advantage on factual knowledge and some agentic tasks. Inkling-Small scores 15.5% on τ³-Banking, compared with 23.7% for Inkling, and its AA Omniscience score is negative, reflecting weaker factual coverage even though its reported hallucination rate is slightly lower.

That tradeoff matters for enterprises. Inkling-Small may be attractive for coding assistants, tool-use systems, retrieval-augmented generation, document analysis and multimodal workflows, but organizations using it for high-stakes factual tasks will still need retrieval, verification and human review.

How a 276B model uses only 12B parameters at a time

Inkling-Small is a sparse Mixture-of-Experts model. According to the model card published by Thinking Machines, its 42-layer decoder routes each token to six of 256 specialized experts, along with two shared experts that remain active for every token.

That architecture helps explain the distinction between the model’s 276 billion total parameters and its 12 billion active parameters. The system retains a large pool of learned capacity but activates only a fraction of it during each inference step.

Advertisement

It is also natively multimodal. Images, audio and text are projected into a shared representation and processed jointly by the decoder rather than being handled through completely separate external systems. Thinking Machines lists coding assistants, agentic applications, chatbots, RAG systems and other multimodal applications among its intended uses.

The company also supports variable reasoning effort, allowing developers to increase or reduce the model’s test-time compute depending on the difficulty of the task. That gives engineering teams a direct way to balance quality, latency and cost across different workloads.

Unfortunately, small does not mean it runs on a laptop

Despite its name, Inkling-Small is not a consumer-scale model.

The standard BF16 checkpoint requires at least 600 GB of aggregate GPU memory, according to Thinking Machines. The company lists two supported configurations: 4x NVIDIA B300 GPUs or 8x NVIDIA H200 GPUs.

Advertisement

A quantized NVFP4 checkpoint lowers the requirement to roughly 180 GB of aggregate VRAM. Thinking Machines says that version can run in W4A4 mode on a single NVIDIA B300, or in W4A16 mode on two H200 GPUs.

That rules out ordinary laptops, MacBooks, desktop gaming PCs and most developer workstations. Even heavily equipped local systems generally fall far short of the required memory.

The practical deployment targets are enterprise GPU servers, cloud clusters and specialized inference providers. The “Small” label is therefore relative to Inkling, not to the broader universe of local models.

Still, the reduction is meaningful. A model that approaches Inkling’s performance while needing substantially less aggregate memory can lower hosting costs, make capacity planning easier and widen the group of organizations capable of self-hosting it.

Advertisement

For companies that want control over data, model behavior and fine-tuning, that smaller footprint may be more important than chasing the highest possible benchmark score.

And of course, it being open source means that it will no doubt be rapidly quantized (made less precise but requiring less compute) and likely blended with other models to be made even smaller for consumer-grade hardware.

Apache 2.0 is the gold standard for enterprise open source models

The licensing may be as important as the benchmarks.

Inkling-Small is released under Apache 2.0, one of the software industry’s most familiar permissive licenses. It generally allows organizations to use, modify, fine-tune, redistribute and commercialize the model, including inside proprietary products, provided they comply with the license’s notice and attribution requirements.

Advertisement

That gives enterprises far more legal flexibility than many custom “open” AI licenses, which may include revenue thresholds, branding obligations, use restrictions or separate conditions for large-scale commercial deployment.

The distinction is increasingly relevant as more AI companies publish model weights without using a conventional open-source license.

Chinese AI darling Moonshot for example, made the weights of its frontier class Kimi K3 model available earlier this week under a custom “open” license that includes additional commercial conditions rather than the comparatively straightforward terms of Apache 2.0.

For legal, procurement and platform teams, that difference can materially simplify adoption. Apache 2.0 does not eliminate the need to review acceptable-use policies, data provenance, regulatory exposure or downstream safety obligations. But it gives organizations a clearer starting point for building internal systems, shipping commercial products and maintaining modified versions of the model.

Advertisement

A more repeatable model-development pipeline

Inkling-Small also shows how quickly Thinking Machines has turned its first large model release into a repeatable engineering process.

Thinking Machines researcher Horace He contrasted the two launches in a post on X:

“Whereas I felt like it took a village to release Inkling, Inkling-Small felt much more routine 😆 We just took the pipeline used for Inkling, passed in a smaller model, and voila — new model! Inkling Small benefited quite a bit vs Inkling from some minor improvements, but there’s still so much more left in the tank…”

The comment suggests the company is no longer treating each model as a one-off research project. Instead, it is building a reusable pipeline for pre-training, post-training, reinforcement learning, evaluation and release.

Thinking Machines says Inkling-Small benefited from an improved pre-training data mix, changes to the machine-learning recipe and on-policy distillation using Inkling as a teacher. The team then continued agentic coding reinforcement learning for two weeks.

Advertisement

Mira Murati emphasized the same point in her own post, describing Inkling-Small as comparable to Inkling at one quarter of the size and highlighting that the weights were open and fine-tunable on Tinker immediately.

How enterprises and AI builders should think about Inkling Small

The company is also distributing full BF16 and NVFP4 checkpoints and supporting deployment through SGLang, vLLM, TokenSpeed, Unsloth and Hugging Face tooling.

That combination gives developers several deployment paths: use an API, fine-tune through Tinker, rely on a third-party inference provider, or operate the model on private infrastructure.

Inkling-Small is not a model that most individuals will download and run locally. But for businesses deciding between a very large flagship and a more manageable open-weight system, it presents a compelling compromise: nearly the same measured intelligence, stronger results on several coding and reasoning tasks, lower token pricing, a smaller hardware footprint and a license that permits broad commercial development.

Advertisement

The broader signal may be just as important. Thinking Machines is showing that Inkling was not a one-time release. The company is already compressing its model family, refining its training pipeline and moving toward a cadence in which open-weight multimodal systems can be produced, improved and deployed more routinely.

Source link

Continue Reading

Tech

New MCP Specification Addresses the Main Barrier To Enterprise Adoption

Published

on

An anonymous reader quotes a report from Ars Technica: This week, the Model Context Protocol (MCP), an open source standard for how AI systems interact with external tools and data sources, saw its largest update since its introduction. Most notably, MCP’s protocol core is now stateless, so requests are no longer dependent on a session tied to an individual server instance. This change has the potential to address long-standing barriers to scalability.

The blog post announcing the specification, written by lead maintainers David Soria Parra and Den Delimarsky (who both work at Anthropic), says: “The highlight of this release is a stateless protocol core — MCP is transforming from a bidirectional stateful protocol into a request/response stateless protocol. It was one of the most highly-requested features from developers who were eager to get better reliability and scalability for their MCP servers.”

[…] There is also a new deprecation policy that ensures at least 12 months between when a feature’s formal deprecation is enacted and when the feature may actually be removed — with a narrow exception for critical security updates. This is again in keeping with the general “let’s make this work better at enterprise scale” theme of the new specification. This update is “MCP’s most important since remote MCP first launched over a year ago,” Soria Parra wrote. Other additions include “Multi Round-Trip Requests, header-based routing, cacheable list results, authorization hardening, a formal extensions framework, and updated Tier 1 SDKs.”

A full list of changes can be found here.

Advertisement

Source link

Continue Reading

Tech

Remember Samsung’s Ballie home robot? It may finally be inching closer to reality

Published

on

Samsung’s rolling home robot has been the industry’s longest-running “will it ever actually ship” joke, and I’ll admit I’d mostly given up on it. Now, there’s finally a glimmer of hope.

Samsung first unveiled Ballie as a concept at CES 2020. Then, it showed it as a significantly upgraded version four years later. However, the company never confirmed a release date or committed to releasing it at all. 

So what did this leak actually reveal?

Now, SammyGuru has obtained the first images of Ballie’s companion smartphone app. To me, it looks like a wireframe design draft rather than a finished product, but it sure isn’t something to ignore. 

The home screen features a status card showing Ballie’s battery level, current position, and any error messages. Then there’s a prominent “Streaming” button alongside it, which appears to let you view the robot’s camera feed and control it remotely from your phone.

The app also includes room-specific shortcuts, letting you send Ballie patrolling through the house or out to greet guests at the front door. One screenshot shows a setup process where Ballie maps your home much as a robot vacuum cleaner does

Advertisement

What can Ballie actually do, and is this really happening?

For context, Ballie is designed to respond to voice commands, control smart home devices, project video onto walls (which is its coolest aspect in my opinion), and function as a mobile security camera when you’re away.

It’s worth tempering expectations here, though. This is still just a wireframe, and there’s no confirmation Samsung is building a functional version of this exact app, let alone a functional version of the smart home robot. But after months of total silence, even a design mockup counts as the most concrete sign yet that Ballie hasn’t quietly died in Samsung’s prototype graveyard.

Ballie’s six-year limbo mirrors the fate of plenty of ambitious CES concepts that quietly vanish without a formal cancellation. The leak suggests internal development is still active, though Samsung’s continued silence on release timing compels me to take this as a hopeful signal.

Advertisement

Source link

Continue Reading

Tech

Tim Cook’s Final Earnings Call: Record iPhone Sales and Future Pricing Woes

Published

on

Demand for Apple’s iPhone and MacBook product lines surged during the third quarter of 2026, signaling record-breaking revenues for June. This was the toplining statement from Apple’s outgoing CEO, Tim Cook, during his final earnings call for the company on Thursday.

This news, which found iPhone sales increasing by 22% to $54 billion, Mac sales of $10 billion (thanks, in large part, to the popularity of its entry level MacBook Neo), the iPad at $6 billion, wearables, home and accessories at almost $8 billion, and services at $31 billion, was relayed with a celebratory tone as Apple raked in roughly $109 billion in net sales for almost $30 billion in net income during the quarter.

“It’s an incredibly strong iPhone and Mac product cycle that has really yielded demand beyond our expectations,” Cook said.

Yet concerns about future price hikes dampened the discussion, given the ongoing global memory shortage and potential supply constraints ahead.

Advertisement

Apple raised prices on many of its products in June due to the ongoing RAM shortage caused by the high memory needs of AI tools and products. As supply chain constraints on memory chips persist, a large portion of the questions asked of Cook centered on price uncertainty, product availability and a potential drop in product sales going forward, depending on economic conditions and customers’ comfort with how much they’ll spend.

“On the pricing front, we reluctantly raised prices,” Cook explained. “I would say we did it because we’re in what I would characterize as a 100-year flood on the memory pricing, with exponential increases in memory prices, and so that was the rationale for it in terms of our philosophy on dollars or percentages.”

What’s the move going forward? Cook said Apple is looking at the bigger picture and thinking about the situation over the long term, rather than treating the next quarter as “a 90-day clock.”

That said, he expects memory pricing to continue its upward trend over the next few months, which means a price hike could be on the horizon. He was cagey in giving any solid details on the matter, which he called “unclear.”

Advertisement

“We’re evaluating all options,” Cook added.

Cook is ending his 15-year run as Apple’s CEO, and before he took questions from members of the press, he took time to express his gratitude, optimism for the company’s future and praise for incoming CEO John Ternus, who will take over on Sept. 1.

“He is truly one of a kind, and there is no better person to take the helm of the company,” Cook said. “I couldn’t be more confident in his leadership, in our executive team and in the extraordinary people at Apple who are determined to enrich the lives of our users all over the world.”

Source link

Continue Reading

Tech

Is the U.S. pulling a ‘China’? Analysts say Washington wants Ukraine’s drone IP without the commitment

Published

on


  • Washington wants Ukraine’s drone tech without signing a formal partnership deal
  • Ukrainian drones now hit targets even after losing operator control entirely
  • Six agreements signed already, ten more contracts reportedly still in the works

Ukraine’s drone program has become one of the most closely watched defense innovations of the current war, drawing global attention.

Analysts now argue that Washington is seeking direct access to that technology’s intellectual property rather than a formal partnership agreement.

Source link

Advertisement
Continue Reading

Tech

EU launches call to build seven AI Gigafactories

Published

on

The Commission hopes to unlock at least €30bn in collective public and private funding.

The EU is pushing forward with its plans to become the “AI continent” with a fresh call for tenders to set up seven AI Gigafactories across the bloc.

The initiative is expected to help the EU develop advanced AI on its own infrastructure in line with the bloc’s rules and standards around model safety.

The Commission today (30 July) said it is setting aside $10bn in EU and national funding, and hopes to tap at least $20bn from the private sector.

Advertisement

Interested consortia or Special Purpose Vehicles (legal entities) have until 12 November to make their bid. Awards will be announced in early 2027, the EU authority said.

The fresh call comes as the EU ramps up efforts to compete with the US in the AI space.

The additional compute capacity is expected to give European enterprises, academia and public authorities access to the infrastructure needed to train and fine-tune advanced AI models. The bloc has an existing network of 19 AI Factories.

Funding is divided into into two consecutive development phases across two lots, with the 18 participating member states, including Ireland, Germany, France and Sweden, matching EU funding, the Commission said.

Advertisement

The first lot of the initiative will support four AI Gigafactories, with each one eligible to receive as much as €100m in EU funding in the first phase, and up to €400m in the second.

A second lot will support three projects, with each receiving up to €200m from the bloc to begin with, and up to €800m per project in a second phase.

Selected AI Gigafactories are expected to begin operations within 18 months after the consortia makes a successful bid.

The EU’s AI Continent action plan, announced in April last year, is a wide-ranging initiative to transform Europe’s strong traditional industries and its talent pool into “powerful engines of AI innovation and acceleration”.

Advertisement

“The opening of the AI Gigafactories call represents a milestone in our AI Continent ambitions. Access to the raw scale of computing power within AI Gigafactories is a strategic necessity for Europe as AI development accelerates,” said EU executive vice-president for tech sovereignty, security and democracy Henna Virkkunen.

“I am pleased to see member states, industry and the Commission coming together to work on these facilities that are key to our technological sovereignty.”

Don’t miss out on the knowledge you need to succeed. Sign up for the Daily Brief, Silicon Republic’s digest of need-to-know sci-tech news.

Advertisement

Source link

Continue Reading

Tech

‘World of Warcraft’ meets ‘Dungeons & Dragons,’ and more from Wizards of the Coast at Gen Con

Published

on

(Wizards of the Coast press image)

This year’s Gen Con is being held in Indianapolis July 30 through Aug. 2. In advance of the show, one of the biggest events on the calendar for North America’s tabletop gaming community, Wizards of the Coast invited press to take an early look at its scheduled announcements for the Dungeons & Dragons roleplaying game.

Wizards, headquartered in Renton, Wash., has published D&D since 1997, and has presided over the game’s meteoric rise in the last decade. Despite renewed competition and several unforced errors, D&D remains the single most popular game in the tabletop roleplaying space, to the point where it’s still often considered synonymous with the hobby as a whole.

In the next year, Wizards’ plans for D&D include a major brand revival, a substantial overhaul to its D&D Beyond online toolset, and a new publishing initiative that promises new and upcoming projects from several of D&D’s classic creators.

What might get the most attention from outside the hobby, however, is a new series of expansions for D&D that will be published under the banner of “Universes Beyond.” The first book for UB, coming on Nov. 17, is a D&D sourcebook based on Blizzard Entertainment’s long-running MMORPG World of Warcraft.

Co-created by long-time D&D designer (and WoW player) James Wyatt, the WoW sourcebook is designed to let players bring the characters, factions, and setting of WoW directly into D&D’s mechanics, with material that spans everything from the earliest days of the MMO to its most recent expansion, Midnight.

Advertisement

Warcraft began in 1994 as a real-time strategy franchise, built around a fantasy kingdom of humans and elves that was abruptly invaded by orcs from another world. The original Warcraft was essentially a fan production for the Warhammer Fantasy universe, but ended up becoming its own original IP.

Warcraft eventually branched out into the nascent MMORPG genre with 2004’s World of Warcraft, which rapidly became a smash hit and is still one of the most popular online games in the current market. WoW has evolved from its original dark fantasy roots into a sort of interplanetary space opera. Its current story arc, the “Worldsoul Saga,” involves the struggle to protect the world of Azeroth from the void priestess who’s out to corrupt it.

The World of Warcraft expansion for Dungeons & Dragons includes rules for playing dracthyr, a draconic species introduced in the WoW expansion Dragonflight. (Wizards of the Coast press image)

This includes D&D versions of trademark WoW species like tauren, dracthyr, earthen, and pandaren; multiple instances from WoW that have been translated into D&D-style dungeon crawls, including a high-level campaign set in Icecrown Citadel; and new D&D subclasses based on character options from WoW, such as death knights, shadow priests, and demon hunters.

This is technically the second time someone’s made Warcraft into a tabletop game, following Sword & Sorcery Studios’ 2003 Warcraft: The Roleplaying Game. According to Wyatt, Wizards’ 2026 book has little to do with that previous publication.

The Warcraft RPG was compatible with 3rd edition D&D but was meant as a standalone product, so it had to spend a lot of its word count on setting up the game’s rules. Wizards’ 2026 World of Warcraft book is an official release for D&D, so it can refer players to the core D&D books for its rules and save its space for more setting details.

Advertisement

The World of Warcraft Universes Beyond book is the first planned entry in what’s intended as an ongoing series of expansions. In the words of D&D franchise head Dan Ayoub, this will extend “Dungeons & Dragons into even more worlds we love.” Exactly which worlds those might be, however, has yet to be announced.

Welcome to Athas, lunch meat. (Wizards of the Coast press image)

Other D&D announcements from Gen Con 2026 include:

One of D&D’s classic settings is coming back in 2027. Dark Sun was originally published in 1991 as a dark, post-apocalyptic take on D&D, taking place on a dying desert world ruled by crazed sorcerer-kings. Wizards’ revival of the setting will be the first official Dark Sun material since 2010, and will be the first Wizards of the Coast product to ship with a mature content warning.

D&D executive producer Greg Bilsland introduced a new initiative he called “Annual Setting Support,” where Wizards will partner with other publishers for yearly new releases to provide additional material for D&D‘s various ongoing campaign settings. 2027 will see a new Ravenloft book made by the Australian company Ghostfire Gaming (Grim Hollow); more Eberron material from Visionary Productions in Canada; and a Forgotten Realms sourcebook by Kirkland, Wash.-based Kobold Press.

Wizards has hired actor and high-profile D&D fan Joe Manganiello (Justice League, Magic Mike) as the creative director for D&D Icons. This new “content program” marks the return of several classic creators to D&D, such as Margaret Weis, Tracy Hickman, and R.A. Salvatore. (Not that Salvatore’s ever exactly left, since he’s been putting out new Drizzt Do’Urden novels almost annually since the 1980s, but you get the idea.)

Advertisement

Luke Gygax is working to finish a project left behind by his late father, Gary Gygax, one of the original co-creators of D&D. The result will be published under the Icons banner, and promises that Luke will “complete an adventure through his father’s eyes.” The result is tentatively planned for publication next year.

2027 will see the release of a new adventure in D&D’s original setting of Greyhawk. Crown of the Witch Queen pits players against the vampire warrior Drelnza, last seen in the classic adventure The Lost Caverns of Tsojcanth, as “she seeks to claim her birthright.”

The forthcoming video game Warlock: Dungeons & Dragons features a vocal performance from Maggie Robertson (Lady Dimitrescu from Resident Evil Village). Robertson voices one of the oldest signature characters in D&D, the archmage Tasha. A first look at Warlock’s gameplay is planned for Aug. 25 as part of the Gamescom conference in Germany.

D&D Beyond plans to revise its mobile experience, which will make it “easier and more intuitive to play D&D your way,” with your phone or tablet as an unobtrusive tool. It can be helpful to keep a browser window open while you play D&D, to look up rules on the fly, and the goal of this new Beyond is to make that process as painless as possible.

Advertisement

Beyond also has a new “Looking for Group” feature in the works, which is designed to help D&D players find people to play the game with, whether it’s locally, virtually, or by hiring professionals to run D&D for them. (The “paid Dungeon Master” has become a genuine thing in recent years. I know. It feels weird to me, too.)

Source link

Continue Reading

Trending

Copyright © 2025