Connect with us

Tech

CXMT replaces Tencent as world’s most valuable Chinese company

Published

on

The manufacturer’s meteoric rise is likely a result of the increased global demand for memory chips.

ChangXin Memory Technologies (CXMT), a Chinese manufacturer of dynamic random-access memory (DRAM) chips, has replaced WeChat creator Tencent as the globe’s most valuable Chinese company. 

According to Bloomberg, as of Thursday (13 August), CXMT has a market capitalisation of $524bn, compared to Tencent’s valuation of $510bn. Established in 2016 in Hefei, CXMT creates the DRAM chips needed to power mobile phones, PCs, tablets, servers, and a range of consumer products and applications. 

Commenting on the announcement, Gary Tan, a portfolio manager at Allspring Global, told Bloomberg, “CXMT exceeding Tencent is a message from the market – chips are the new clicks. Our sense is the gap between the two will widen as agentic AI takes an increasing share of internet flows.”

Advertisement

CXMT’s overtaking of rival Tencent comes at a time when there is a major global push by organisations to invest in the development of chips, alongside increased interest in tech and resource sovereignty. 

Reportedly, the chip manufacturer is representative of the direction many Chinese organisations are leaning towards in regards to AI and semiconductor targets. Shares are currently up a further 8pc following CXMT’s initial 467pc surge in debut figures. 

Globally, organisations are seeking to mitigate the chips shortage, outpace their rivals and secure a more reliable supply chain in order to advance their products, tools and technologies. 

This month, AI giant Anthropic confirmed plans to design its own chips in response to the worldwide shortage and increased pressures to develop faster, more advanced AI systems. 

Advertisement

In July, semiconductor and infrastructure manufacturer Broadcom extended the partnership it holds with Apple, in a deal that will see both organisations collaborate on custom-made chips until 2031. Apple is also currently working on AI server chips. 

Don’t miss out on the knowledge you need to succeed. Sign up for the Daily Brief, Silicon Republic’s digest of need-to-know sci-tech news.

Source link

Advertisement
Continue Reading
Click to comment

You must be logged in to post a comment Login

Leave a Reply

Tech

Writer says its new Palmyra X6 model cuts AI agent costs by 52% as token spending surges

Published

on

Writer, the enterprise AI agent platform used by Fortune 500 companies including Accenture, Uber, and Vanguard, released its new flagship model Palmyra X6 today, alongside a rebuilt agent orchestration “harness” and new governance tools designed to give IT leaders control over runaway token spending.

The headline numbers are striking: Writer says its agent product now operates at an average 52% lower cost, with a 48% improvement in speed and a 10% improvement in quality when paired with Palmyra X6. But the more consequential story may be how the company got there — and what its choices reveal about where the enterprise AI market is heading.

Palmyra X6 is not trained from scratch. It is a post-trained version of GLM-5.2, the open-weight mixture-of-experts model from Beijing-based Z.ai, formerly Zhipu AI — a fact Writer discloses openly in its technical report, and one that places the San Francisco company at the center of one of the industry’s most charged debates: whether American enterprises should build on Chinese open-source foundations.

“This model is in no way, shape, or form connected to any of its original developers. It is fully run on our U.S. infrastructure,” Matan-Paul Shetrit, Writer’s director of product management, told VentureBeat in an exclusive interview ahead of the announcement.

Advertisement

Dan Bikel, who leads Writer’s AI research, put it more bluntly: “It’s very much a Palmyra model, and we just happen to grab the floating point numbers as the starting point, and train from there.”

Why AI agents are blowing up enterprise budgets in ways chatbots never did

Writer’s announcement lands at a moment when the economics of agentic AI have moved to the center of enterprise buying decisions. Unlike a chatbot, which typically generates one answer per user request, an AI agent turns a single request into repeated rounds of planning, retrieval, tool calls, validation, and retries — with every loop consuming metered tokens. The user sees one answer; the invoice reflects the entire loop.

The scale of the problem is becoming clear. Goldman Sachs forecasts that token consumption will multiply 24 times between 2026 and 2030, reaching 120 quadrillion tokens per month, driven not by more people asking questions but by always-on enterprise agents. The same analysis warned that falling per-token prices do not guarantee falling bills: if an agentic task draws 20 times more tokens while unit prices fall 75%, total charges still rise fivefold.

“The enterprise wants token consumption to explode — it means adoption is happening — but they need costs to flatten,” said Waseem AlShikh, Writer’s CTO and co-founder, in a statement.

Advertisement

Shetrit framed the cost problem as the primary obstacle to enterprise AI adoption — more so than model capability itself. “The biggest barrier today to enterprise expansion using AI is actually not model capabilities in most cases; it’s actually the cost around them,” he said. “The reality today is, in most cases, the alternative for AI is not another AI, it is human labor.”

Asked whether cutting customers’ token consumption would cannibalize Writer’s own per-token revenue, Shetrit rejected the premise. “Reducing the cost is not hurting my bottom line. It’s actually expanding it, because it’s expanding the TAM of opportunity within an organization,” he said, arguing that lower per-task costs unlock workflows enterprises would otherwise never automate. That argument echoes a pattern familiar from the cloud era, where unit prices fell for a decade while total bills rose as consumption expanded — a dynamic Writer is explicitly betting will repeat with agents, and betting it can profit from.

Inside Palmyra X6: how 626 training examples fine-tuned a 744-billion-parameter model

Palmyra X6 is a 744-billion-parameter mixture-of-experts model with roughly 40 billion active parameters per token, inheriting GLM-5.2’s architecture unchanged, according to Writer’s technical report. The company’s contribution is a deliberately conservative post-training recipe: a technique called anchored supervised fine-tuning (ASFT), applied to a remarkably small corpus of just 626 curated synthetic agentic trajectories, trained for a single epoch at a low learning rate.

The tiny dataset is the point, not a limitation. ASFT pairs a token-weighting scheme with a KL-divergence “anchor” that penalizes the fine-tuned model for drifting too far from a frozen copy of the base model — teaching new tool-use behaviors without eroding the general capabilities the base already has. Writer also swapped the standard Adam optimizer for Muon, a newer method that treats weight matrices as geometric objects, on the model’s core weight matrices.

Advertisement

“There’s a whole string of papers following a quote-unquote ‘less is more’” philosophy, Bikel said, referencing research showing that “small, extremely high quality data sets go a really long way.” He added: “That’s the philosophy — one of the philosophies — that we followed when building this model, and it showed. It allowed us to optimize for our customers at lower cost to do the work of optimization, and that ultimately yielded a lower cost model for us and for them.”

The training data itself is fully synthetic — every plan, tool call, and final answer machine-generated by teacher models, then filtered through structural quality gates, a model-based verifier, and a two-model LLM judging panel before entering training. That continues a long-standing Writer practice: the company’s Palmyra X 004 was trained almost entirely on synthetic data for roughly $700,000 back in 2024, as TechCrunch reporte at the time, and Palmyra X5 required about $1 million in GPU hours, according to SiliconANGLE.

On Writer’s internal evaluations — nine capabilities spanning grounding and retrieval, tool use, content generation, sub-agent delegation, and brand voice — X6 scored an average of 0.87 out of 1.00, edging out Anthropic’s Claude Opus 4.8 (0.86), Claude Sonnet 4.6 (0.85), OpenAI’s GPT-5.5 (0.80), and Google’s Gemini 3.1 (0.77). The price gap is the real differentiator: Writer prices X6 at $2 per million input tokens and $8 per million output tokens, versus 15/75 for Opus 4.8. The company says X6 completes tasks in 26 seconds on average and can work unattended toward a single goal for up to eight hours.

Writer is candid that internal benchmarks invite skepticism. Asked directly whether the company would publish its methodology after grading its own homework, Bikel said the technical report covers “both the protocol we used to do our public benchmarking as well as our internal evaluations.” He described public benchmarks as sanity checks rather than targets: “We do things like public benchmarks to let us know that we’re climbing the right hill and that we don’t have any sort of huge gaps, but we don’t slavishly follow them either, because that’s not really serving our customers.”

Advertisement

The China question: what building on GLM-5.2 means for enterprise security and trust

Writer’s choice of base model would have been unthinkable for an American enterprise vendor two years ago. Today it reflects a market reality: GLM-5.2, released in June under the permissive MIT license, is arguably the most capable openly available model in the world. Independent analysis house Artificial Analysis scored it at 51 on its Intelligence Index — ahead of DeepSeek V4 Pro, Kimi K2.6, and even some of Google’s Gemini models on agentic tasks — while undercutting U.S. flagship API pricing many times over, as European tech outlet Trending Topics reported. Writer’s press release calls it “the strongest available open-weight model.”

The open-weight surge carries genuine baggage. An August report from AI safety nonprofit SaferAI found that GLM-5.2 refused none of the offensive cyber or biology tasks it was given via Z.ai’s public API, and that Z.ai published no safety framework or pre-deployment risk assessment — a gap that widens once anyone can download and modify the weights.

Writer’s answer is that provenance and post-training matter more than origin. Bikel emphasized that the company “grabbed the weights off of the U.S. Hugging Face” and trained entirely on American infrastructure; the technical report states all datasets were synthesized and stored in the U.S., and all training hardware was located in the U.S.

The company also ran what it describes as an unusually rigorous, pre-registered model-risk evaluation covering political bias, censorship, factuality, and refusal behavior — 19,674 evaluated responses scored by blinded judges — comparing X6 against its GLM-5.2 base and four frontier control models.

Advertisement

On the Washington Post’s ModelSlant political-bias evaluation, Writer says X6 presented both sides of hot-button questions 80% of the time, the highest rate of any model tested, and answered politically sensitive prompts that DeepSeek V4 refused outright. On the FORTRESS adversarial safety benchmark, X6 with its deployment system message scored 8.6 points higher on adversarial safety than the raw GLM-5.2 base, at negligible cost to benign helpfulness.

“We’ve run extensive benchmarking around bias, around censorship,” Shetrit said, “and the work Dan and the team has done has actually proven that this model is actually significantly better than not just open source alternatives, but any closed source alternative in the market at the time of the benchmarking.” 

The report does hedge in one notable place: while English-language behavior showed no statistically robust political asymmetry, “the behavior was shown to vary by language” — a candid admission that 626 fine-tuning trajectories do not scrub every trace of a base model’s training.

The harness effect: why orchestration may matter more than the model itself

Perhaps the most strategically interesting claim in Writer’s announcement has nothing to do with Palmyra X6 at all. The company says its rebuilt Writer Agent harness — the orchestration layer that plans tasks, batches work, delegates to sub-agents, and manages context — cuts costs by 41% and completes tasks 44% faster across every model it tested, including third-party models from Anthropic and OpenAI, while maintaining quality. Writer published the finding in an accompanying research paper on what it calls “The Harness Effect.”

Advertisement

That raises an obvious question, which VentureBeat put to the company: if the harness alone delivers most of the savings on any model, why build a model at all?

Shetrit’s answer was about control. “I cannot control if a lab deprecates their model. I cannot control what data they use in their model,” he said. “Where when I build the model, I have significant moral control, and I can answer the tough questions that enterprise customers ask me.”

Bikel added that the model and harness were developed together: “This model was built and essentially co-evolved with the harness… We know that we have a flagship product, Writer Agent. We want that to work really, really well with this model, and sure enough, it does. And we take that into account during model development, and that’s something that is not possible if you don’t build your own model.”

Notably, Writer is simultaneously hedging. With this release, the company extends multi-model support to Writer Agent, letting admins enable models from Anthropic, OpenAI, and cloud providers including Microsoft Azure, AWS Bedrock, and Nvidia NIM — even image-generation models, a category Writer does not build. The message to CIOs is disarmingly simple: use our model because it is cheapest and best for your workflows, but the platform saves you money either way.

Advertisement

New governance tools aim to end surprise AI bills before they start

The third leg of the release targets a quieter enterprise pain point: nobody in the C-suite knows what the agents are spending. New governance tools give administrators a centralized view of agent usage across the business, per-workflow analytics for the company’s shareable “Playbooks” and “Skills” automations, and consumption controls with alerts and spending limits.

Asked whether the introduction of spending controls implied that customers had been receiving surprise bills, Shetrit reframed it as an adoption enabler rather than damage control. “How do we build the tools to allow you as the CIO, CISO in a company, to feel comfortable both on the security and spend, so you can expand AI usage in your organization,” he said. In his telling, visibility is what lets leaders say yes: businesses with clear cost data “are actually looking to expand AI adoption to use cases that they would never have touched before.”

The feature set tracks a broader shift in how enterprises budget for AI. As Forbes analysis of the token price wars argued, sophisticated buyers are learning to model cost per successful task — counting retries, tool calls, and escalations — rather than multiplying expected calls by the advertised rate card. Writer is effectively productizing that discipline, turning what has been a finance-team spreadsheet exercise into a native platform capability.

It also completes a governance arc the company has been building for over a year. Writer shipped its unified agent experience with admin controls last November, then added agent Skills and workflow analytics in March, according to earlier company announcements. Thursday’s release closes the loop by attaching a price tag — and a spending limit — to every workflow.

Advertisement

Writer, founded in 2020 by May Habib and Waseem AlShikh, raised $200 million at a $1.9 billion valuation in late 2024, and has built its business on regulated, high-stakes deployments rather than consumer scale. Shetrit made no apology for the narrowness of that focus. “The privilege of working and focusing on enterprise use cases is that I don’t need my model to be able to write a French sonnet,” he said. “When you don’t try to do everything, you can focus on your customer problem and needs.”

He was equally direct about identity: “We are not a research lab converted to a consumer product now dabbling in enterprise. We are first and foremost an enterprise company that serves enterprise customers, and we evaluate our decisions within that lens. Which means, if we think building things from scratch is the right decision, that’s what we will do. But if we think there are other alternatives out there in the market that serve our customers better, that’s what we will do.”

That pragmatism may be the release’s most important signal. A well-capitalized American AI company with five years of model-building experience has concluded that the frontier of value no longer lies in pretraining, but in the last mile: post-training open weights, engineering the harness around them, and handing the CFO a dashboard. If Writer is right, the frontier labs’ moat narrows to the workloads where quality genuinely justifies a sevenfold price premium — and for everything else, the winning model is the one somebody else paid to pretrain.

In an industry that has spent three years arguing about whose model is smartest, Writer is making a different wager: the enterprise AI race won’t be won by the company with the best floating point numbers, but by the one that knows what to do with them.

Advertisement

Source link

Continue Reading

Tech

Google’s cheap model is now two versions ahead of its flagship

Published

on

Google has released Gemini 3.7 Flash with sharp gains on coding benchmarks and introductory pricing of $0.75 per million input tokens. Gemini 3.5 Pro remains months behind schedule, and Google will not say whether it is still coming.

Google has released Gemini 3.7 Flash, and still will not say when its flagship model is coming. The workhorse model posts large gains on coding, which is the ground Google has been losing. Gemini 3.5 Pro, the model this one is meant to sit beneath, remains months behind schedule.

The numbers are real. On FrontierCode 1.1 the new model scores 43.6% against 34.4% for its predecessor, on DeepSWE v1.1 it reaches 65.3% against 49.0%, and on AutomationBench it more than doubles, from 17.0% to 30.4%.

Google is pricing it to be used. Input costs $0.75 per million tokens and output $3.75 until the end of December, after which both double. It also powers Gemini Spark for AI Pro and Ultra subscribers across more than 160 countries.

Advertisement

The safety work is pitched at the frontier categories. Google says it added safeguards against cyber offence and chemical, biological, radiological and nuclear misuse, in line with its Frontier Safety commitments.

What stands out is the cadence. This lands three weeks after the last Flash release, and Google put out three Flash models in July alone, including a security-tuned variant.

The Pro line has not moved in that time. Google’s next Pro model is months behind schedule after its coding performance fell short of internal targets, which is awkward given coding is where the money is.

It may not arrive at all. Axios reports Google could skip 3.5 Pro entirely and go straight to Gemini 4 Pro, and the company declined to discuss the model’s fate.

Advertisement

Sundar Pichai told investors in July that Google intends to ship models faster and is already spending heavily on compute to train Gemini 4. Shipping the cheap tier every three weeks is one way to look faster while the expensive one stalls.

There is a cost inside the building. The delay has been hard on DeepMind morale, at a moment when rival labs are hiring.

Source link

Advertisement
Continue Reading

Tech

Honor’s Robot Phone Officially Launches, Moves Its Camera So You Don’t Have To

Published

on

Honor Robot Phone Launch
Honor put a motorized camera arm on a regular-looking phone and started selling it in China yesterday. The device, simply called the Robot Phone, costs 9,999 yuan for the 12GB and 512GB version, roughly $1,500, with a higher 16GB and 1TB model at 12,999 yuan. Pre-orders opened immediately after the Guangzhou event, and full sales begin August 18. For now the phone stays in China.



When first turned on, the phone appears to be your standard device. Its dimensions are normal to say the least, measuring 9.59 mm thick, weighing 248 grams, and having a sleek metal body with modest, curved corners. There’s a 6.31-inch LTPO OLED display up front, which can produce 6,800 nits and switch between 1 and 120 frames per second seamlessly. All of this is powered by a Snapdragon 8 Elite Gen 5 chip, which also includes a 7,060 mAh battery, rapid charging (120 watts wired, 50 watts wireless), and just IP54 water protection.

Sale


Google Pixel 10a – 30+ Hours Battery, Camera Coach, Gemini – Fog 128GB
  • Google Pixel 10a is a durable, everyday phone with more[1]; snap brilliant photography on a simple, powerful camera, get 30+ hours out of a full…
  • Unlocked Android phone gives you the flexibility to change carriers and choose your own data plan; it works with Google Fi, Verizon, T-Mobile, AT&T…
  • Pixel 10a is sleek and durable, with a super smooth finish, scratch-resistant Corning Gorilla Glass 7i display, and IP68 water and dust protection[4]

When you press a button, or even just wave your hand, something interesting happens. A glass cover swings open, showing a small titanium gimbal protruding from the rear of the phone. Someone at Honor evidently spent a long time figuring out how to make this thing smaller, as the mechanism is 65% shorter than typical, yet it also contains over 100 precise pieces. The 2.6-gram titanium motor can spin at a staggering 360 degrees per second. Want the camera to remain in a fixed position? Or keep your first-person perspective? Perhaps you just want to spin it 90 or 180 degrees on a whim? Not a problem; the arm can do it all, and more. Once up and going, the camera maintains its level even as the rest of the phone jerks around.

Advertisement


The main sensor is also impressive, at 200 megapixels and powered by a 1/1.28-inch chip and f/1.6 lens. You get a matched 200 megapixel telephoto with 2.7 times optical zoom, as well as a 50 megapixel ultrawide lens capable of shooting at angles of up to 122 degrees. Honor has partnered with ARRI to enable the phone to take footage in 10-bit LogC3, using the official ARRI Wide Gamut and all that jazz. When you take a photo, it is processed by Honor’s H1 imaging technology, which is excellent at reducing noise and other issues. Do you want to record some serious high-speed footage? There’s no need to worry because 4K at 120 frames per second is equivalent to 1080p at 240 fps. The cherry on top is that you can import the clips directly into DaVinci Resolve with the ARRI LUTs already applied.


The software portion is where things get very interesting. When you press a few buttons or tap a couple of faces, the YOYO Robot Mode activates, autonomously following individuals and keeping the camera focused on them without your interaction. Double tap a face to lock in, even if they stroll away for a bit, and it will immediately reclaim them. Alternatively, use gesture controls to flip the arm up and start rolling.


Want to tell the camera what to do? Voice instructions are available on the menu, and you can instruct it to swap angles or turn for a specific number of seconds. Still, all of this debate about modes and capabilities misses the point: this phone can be used to create video that is nearly identical to the original. Take a solo stroll or go live on a treadmill, and the phone will track the movement with much less tremor than if you held it by hand. Group photographs are also a breeze, since the arm swings from side to side, capturing a few shots and quickly stitching them together to create a lovely, broad panorama.
[Source]

Advertisement

Source link

Continue Reading

Tech

Mystery attacker spent a year raiding Salesforce and ServiceNow portals

Published

on

CYBER-CRIME

Custom tools harvested whatever over-permissioned guest accounts would surrender

Someone has spent more than a year rifling through Salesforce and ServiceNow portals around the world, harvesting data that organizations accidentally left open to anyone who came looking.

Researchers at Reco have named the operation “City-Forum” after a domain connected to its infrastructure. The domain has pointed to the attacker’s server since March 2025, although exactly when the campaign began is unclear. Reco says the activity is continuing and increasing in volume.

Advertisement

Reco isn’t naming the targets, but said it spotted the attacker poking around portals belonging to telecoms companies, banks and other financial services firms, enterprise software vendors, cybersecurity companies, and public sector bodies.

“In the last year, we’ve seen many threat actors that use Aura enumeration against over-permissioned Salesforce guest users. This actor is different,” said Nitay Bachrach, senior security researcher at Reco.

On Salesforce, the attacker targets Lightning Web Runtime (LWR) sites through the UI API’s GraphQL layer, an approach Reco says it has not found documented in public research or incorporated into publicly available attack tools. Over at ServiceNow, the same operator queries a native Service Portal search endpoint that has received little public attention.

If the guest can read a record, so can anyone on the internet. That is not a platform vulnerability

Security researcher Nitay Bachrach

Advertisement

The tooling also checks whether Salesforce sites permit self-registration, potentially offering a route from anonymous guest access to an authenticated external account with permission to see considerably more data. Reco said it saw these checks across most of the Salesforce targets it examined.

“The threat actor created their own toolset, based on research and techniques which are not well documented online,” Bachrach said. “They studied the services to map different common data leak vectors – this is an advanced actor.”

This isn’t casual poking around either. Reco said the busiest Salesforce target logged more than 560,000 events from the attacker’s IP during the campaign, almost all attempts to enumerate data available to guest users.

Advertisement

Reco linked the Salesforce and ServiceNow activity to the same server, which targeted multiple organizations around the world. More unusually, the attacker hasn’t bothered changing its infrastructure: the same IP address and domain have remained in use for at least 17 months, with related custom tooling doing the rounds across both platforms.

ServiceNow told us it is “aware of a security company’s blog post claiming certain configurations are creating security risk. As noted in the security company’s post, there are no allegations of a compromise of the ServiceNow environment. Nonetheless, we take third party reports seriously and are investigating accordingly. Our priority is to protect our customers, their data, and our systems.”

Salesforce has not yet responded to The Register‘s questions. 

Salesforce customers have already had one very public lesson in what can happen when guest access gets too generous. In March, ShinyHunters told The Register it had stolen data from around 100 high-profile companies and nearly 400 websites after going after over-permissioned Experience Cloud guest accounts.

Advertisement

City-Forum isn’t doing quite the same thing, and Reco isn’t blaming ShinyHunters. “We don’t know who this is, and we’re not ruling anyone in or out,” Bachrach said.

Reco says all the activity it observed was conducted without authentication, with the attacker collecting information that organizations had exposed through permissions, sharing rules, search sources, or other configuration choices.

“If the guest can read a record, so can anyone on the internet,” Bachrach warned. “That is not a platform vulnerability.”

Which is good news for Salesforce and ServiceNow, perhaps, but rather less comforting for anyone now wondering what their guest account has been showing the guests. ®

Advertisement

Source link

Continue Reading

Tech

Why Capital One built its multi-agent AI platform around open-weight models

Published

on

Presented by Capital One


At VB Transform 2026, Kel Vanee, MVP of machine learning engineering at Capital One, spoke with Sam Witteveen, Senior Technology Contributor at VentureBeat, about how the bank built a scalable multi-agent AI architecture around deeply customized open-weight models rather than relying on an off-the-shelf foundation model.

“At Capital One, we’re not just using AI, we’re building AI,” Vanee said.

Advertisement

The groundwork was laid years ago with Capital One’s early investments in data transformation and cloud adoption, which Vanee said were foundational to moving quickly when the current wave of AI arrived. That technical foundation enabled the company to make several deliberate architectural decisions, including building a centralized, enterprise-wide AI platform with built-in governance, deeply customizing open models with proprietary data, and constructing its own multi-agent orchestration harness.

Customizing open-weight models with proprietary data

Rather than relying solely on off-the-shelf frontier models, Capital One fine-tunes open-weight models using its rich, proprietary data.

“We view our data as a huge advantage and something that nobody else has, something that the general frontier models cannot provide. So we are taking that data and deeply customizing these models,” Vanee explained. He added that real-time data is absolutely critical to bring in fresh context during live customer or associate interactions.

Vanee also revealed an unexpected benefit of this approach: extensibility across the enterprise.

Advertisement

“As we customize those open-source models for one use case, we actually see benefits across our whole portfolio,” he noted. “We are training that model to be an expert at Capital One use cases, policy, and nomenclature. As we do that training, we see a general lift.”

Inside Capital One’s multi-agentic AI workflow

As an example of the approach, Vanee pointed to a customer-service workflow for bank fraud that handles millions of calls a year, where interactions range from roughly four minutes to as long as sixty minutes, and where an initial attempt at engaging a single large language model proved insufficient. With Capital One’s multi-agentic workflow (MACAW), interactions are routed through specialized agents with governance and guardrails built in.

“The MACAW workflow is made up of a number of different agents,” he said. “The first one is an understanding agent. Its purpose is to look at what the customer is saying and try to understand what their intention is.”

From there, a reasoning agent is given several specific instructions to generate a summary; a validation agent fact-checks the summary to ensure it is accurate; and an explaining agent turns the summary into a formatted document with all necessary details that is then shared with agents.

Advertisement

For the consumer banking use case, this workflow helps several hundred customer-service agents who specialize in complex fraud calls. The post-call summaries it generates help document long, back-and-forth interactions that agents previously had to reconstruct by hand.

Capital One’s multi-agentic architecture also underpins Chat Concierge, a customer-facing auto-shopping assistant, which further leverages a version of Meta’s open-weight Llama model that has been customized with Capital One’s proprietary data. It uses the same division of labor, with one agent conversing with the customer, one building an action plan from business rules, one evaluating accuracy, and one explaining and validating the result.

Optimizing latency and cost with an agentic research system

Beyond customer-facing solutions, Capital One is also leveraging agentic AI to automate rote tasks for its employees and help them focus on high-leverage aspects of their work. In one example, the company built an autonomous agentic optimization solution to tune backend hosting infrastructure.

Vanee explained that in the world of LLMs, where new optimizations are delivered every day, they aren’t all complementary. Combining two good optimizations can sometimes cause a performance regression.

Advertisement

“This agentic system will run through a search space that is designed by the researcher, handle all the mechanics of setting up that experiment and running the experiment, and then put a whole summarization of the results in front of the researcher,” Vanee said.

Vanee added that the system allows researchers to “find the series of optimizations and configurations that’s really going to give [them] the best latency possible.”

What’s next: model routing and proactive, event-driven AI

Looking ahead, one big trend Vanee sees is routing abstraction layers that a platform seeks to validate over multiple models, both for cost and accuracy.

“We actually think that you can get better accuracy than any individual model simply by routing across a broader set of available models, because different models are going to excel in different areas,” he said.

Advertisement

His second prediction was a shift toward systems that act without waiting to be asked, while also emphasizing that deploying such proactive agents would demand rigorous testing and monitoring.

“The thing I think is going to become bigger in the future is more proactive and event-driven AI,” Vanee said. Rather than waiting for a human prompt, AI would step in as soon as it detects conditions that warrant action.

“This is going to enable more monitoring and larger-scale monitoring, and it’ll empower us as we fight fraud and address these opportunities,” Vanee said. “So proactive AI is going to be a really important trend.”

Driving continuous AI innovation in financial services

Capital One’s approach underscores a broader truth for enterprise technology leaders: driving measurable value with AI requires moving beyond off-the-shelf software toward deeply customized, highly governed architectures. By combining fine-tuned open-weight models, a multi-agent orchestration harness, and proprietary data assets, the bank has established a repeatable blueprint for deploying scalable AI in financial services.

Advertisement

“All of those ingredients were absolutely critical to differentiating in this space and hitting the quality bars as well as the cost and latency thresholds we set for ourselves,” Vanee said.

As the company expands these capabilities across new use cases, its enterprise platform approach helps to ensure that technical breakthroughs translate into safer, faster, and more personalized experiences for its millions of customers.


Sponsored articles are content produced by a company that is either paying for the post or has a business relationship with VentureBeat, and they’re always clearly marked. For more information, contact sales@venturebeat.com.

Source link

Advertisement
Continue Reading

Tech

IBM, OpenAI join forces to scale AI adoption and boost security

Published

on

Specialised engineers and consultants trained through OpenAI’s partner network will work directly with clients on solutions.

OpenAI – in its latest effort to garner enterprise support – has teamed up with IBM to provide organisations with its tools to scale AI deployment.

The deal will see the two companies create industry-specific, go-to-market initiatives targeted at financial services, government, telecommunications and retail sectors, and comes at a time when AI investments are surging multi-fold despite only a fraction of businesses seeing real returns.

“While enterprises are rapidly investing in AI, they are looking for practical ways to apply it across their core operations to deliver measurable business outcomes and create new commercial models,” said Andy Baldwin, the global senior vice-president at IBM Consulting.

Advertisement

“The challenge is not access to AI technologies – it’s integrating AI securely and at scale into complex enterprise environments and workflows.”

The partnership embeds OpenAI’s latest frontier models, including GPT-5.6 and products like Codex and ChatGPT Work, into IBM’s AI platform.

IBM said it will deploy units of specialised engineers and consultants trained through OpenAI’s partner network to work directly with clients to accelerate AI implementation across complex business workflows and highly regulated environments.

The technology giant is also launching a dedicated channel through which more of its consultants and engineers can be certified under the partner network.

Advertisement

The two also said they will expand their collaboration as part of OpenAI’s cybersecurity partnership programme Daybreak by combining the AI giant’s frontier capabilities with IBM’s multi-agent-powered service for delivering decision-making and intelligence.

With this, the collaborators aim to help clients manage both cyber and AI model risk, including application-layer vulnerabilities, governance gaps and operational risks that limits AI adoption, they said.

“The organisations pulling ahead with AI are the ones turning it into a trusted part of how their business operates,” said Denise Dresser, the chief revenue officer at OpenAI.

Earlier this year, IBM set aside more than $10bn for its lofty quantum plans that it hopes will deliver the world’s first large-scale, fault-tolerant quantum computer by 2029.

Advertisement

Don’t miss out on the knowledge you need to succeed. Sign up for the Daily Brief, Silicon Republic’s digest of need-to-know sci-tech news.

Source link

Advertisement
Continue Reading

Tech

Is it the right time?

Published

on

Verdict

The Samsung 990 is a solid PCIe 4.0 SSD with middle-of-the-pack sequential and random performance that’s decently brisk, alongside sensible capacity options and decent durability. For a more value-oriented drive, though, its retail pricing is some of the most baffling I’ve seen.

  • Solid speeds

  • Single-sided design is handy for PS5 use

  • Speedy file transfer rates

  • Heinously expensive

  • Performance not as strong as rivals

SQUIRREL_PLAYLIST_10208689

Advertisement

Key Features

  • PCIe 4.0 Standard:

    The 990 is a PCIe 4.0 SSD, meaning you’re getting blazing fast speeds for its standard. It may not be as quick as a Gen 5 option, but this one is using as much of the Gen 4 standard as it can.

    Advertisement
  • Up to 2TB capacity:.

    It also comes in reasonable capacities up to 2TB to make it a solid choice for storing a good range of games and apps on

    Advertisement
  • PC and PS5 compatible:

    The 990 will play nicely in a compatible M.2 slot for PC use, while its speeds also make it compatible with PS5, as long as you grab an inexpensive heatsink.

    Advertisement

Introduction

In its infinite wisdom, Samsung has decided now is the ideal time to release an ‘affordable’ PCIe 4.0 SSD with its 990 drive.

That’s a bit of an odd decision with the high pricing we’re seeing for memory and storage given the current circumstances, but I nonetheless appreciate its efforts. The new 990 is, to all intents and purposes, a QLC-based variant of its PCIe 5.0 990 Evo Plus drive, and aims to bring another option for folks wanting a cheaper PCIe 4.0 SSD for use in a PC or PS5.

There is some stiff competition that Samsung has to go up against, with the likes of the WD Black SN7100 and the Crucial P310 both key rivals, and the brand has to do it with a drive that’s $269.99 for the 1TB variant I have here. Even with the current circumstances, that pricing feels baffling.

Nonetheless, I’ve been using the 990 in my PC for the last couple of weeks to see if it’s one of the best SSDs we’ve tested. Let’s find out.

Advertisement

Advertisement

Specs

  • Solid storage options on offer
  • No DRAM cache as HMB is used instead
  • Reasonable endurance rating and warranty

The 990 isn’t overly exciting by way of looks, with an all-black frame and a small sticker on the front showing the Samsung logo and model designation over the NAND chips.

It is a single-sided drive for better thermals under a heatsink, and this is a drive that’s compatible with both PC and PS5, being both the standard M.2 2280 size and form factor that’ll fit in the respective PCIe 4.0 slot. 

Rear - Samsung 990Rear - Samsung 990
Image Credit (Trusted Reviews)

The sample I have is a 1TB option, although it is also possible to get the 990 in a larger 2TB model if you want more storage.

Advertisement

The controller that the 990 features is Samsung’s own PiccoloQ controller – the same as in the 990 Evo Plus – while the NAND is also Samsung’s own V9 QLC. You won’t find a DRAM cache here, either, as Samsung has opted to go for HMB, the same as with the WD Black SN7100.

Advertisement
Profile - Samsung 990Profile - Samsung 990
Image Credit (Trusted Reviews)

This isn’t uncommon for more affordable drives and isn’t a big issue for most use cases outside of intensive scientific computing, instead using system RAM for caching data as opposed to a dedicated cache on the drive itself.

One thing that is strange is that the speeds of the drive aren’t consistent across the range. The 2TB drive is slightly faster in terms of its sequential reads as well as its random read and write speeds. It’s because the 2TB variant has enough flash memory working in parallel to make the most of the controller and to get more optimal performance.

Profile - Samsung 990Profile - Samsung 990
Image Credit (Trusted Reviews)

The reliability here is scaled across the two capacities, with Samsung offering a 400TBW for the 1TB model and 800TBW for the 2TB variant. This is in line with key rivals, although isn’t as spectacular as other drives we’ve tested, such as the Kingston Fury Renegade‘s 2000TBW rating.

Advertisement

Full Specs

  Samsung 990 Crucial T500 WD Black SN7100 Crucial P310
Connector M.2-2280 M.2-2280 M.2-2280 M.2-2280
Interface PCIe 4.0 x4 PCIe 4.0 x4 PCIe 4.0 x4 PCIe 4.0 x4
Model Variants 1TB, 2TB 500GB, 1TB, 2TB 500GB, 1TB, 2TB 500GB, 1TB, 2TB
Read Speed 7250MB/s 7400 MB/s 7250 MB/s 7100 MB/s
Release Date 2026 2023 2024 2024
Storage Capacity (Sample) 1TB 2TB 1TB 2TB
USA RRP (2TB) $529.99 $149.99 $149.99 $137.99
Write Speed 6450 MB/s 7000 MB/s 6900 MB/s 6000 MB/s

Test Setup

Of course, for testing any quantity of PC components, SSDs included, I needed to make sure I had a solid PC to do so. Hence, I took the decision back in early 2024 to upgrade my ailing HP pre-built to a fully custom rig with a system that benefits from brisk gaming performance and excellent compatibility with modern and future hardware.

The full system specs can be found below:

Advertisement
  • CPU: AMD Ryzen 7 7800X3D
  • Motherboard: NZXT N7 B650E
  • GPU: Nvidia RTX 4080 Super Founder’s Edition
  • RAM: 32GB (2x16GB) Corsair Vengeance DDR5-6000 CL36
  • Cooler: Noctua NH-D15
  • PSU: 1200W NZXT C1200 80+ Gold ATX 3.0
  • Case: NZXT H9 Flow

The long and short of the setup is that the Samsung 990 was placed in a compatible PCIe 4.0 x4 slot on my B650E motherboard, and then a range of real-world and synthetic tests were run. These included the classic CrystalDiskMark 8 with its Sequential speeds at a queue depth of 8 and 1, as well as its Random 4K performance at depths of Q32 and Q1. The Sequential tests are handy in proving the actual raw speed of the drive for fast file copies and access, while the Random 4K tests are more indicative of loading a game up.

For the usefulness of a quantifiable ranking, I’ve also included the Quick System Drive and Data Drive benchmarks from the PCMark 10 suite.

Advertisement

As for real-world testing, I’ve elected to see the transfer rates in moving over a set of test files totalling 120GB (in reality, a set of ripped Blu-Rays of recent Marillion concert film and hi-res audio) using the Windows File Explorer, noting down its average transfer rate, and to see the speeds at which it can move over the 110GB Dirt Rally 2.0 using Steam. As for game loading times, I’ve taken note of how quickly the Optimus GX7100M runs the Final Fantasy XIV Endwalker standalone benchmark running at 1080p and Maximum settings, as is consistent with our other testing.

Performance

  • Decent sequential performance
  • Random performance in the ballpark against rivals
  • PCMark10 results are a little off

The 990 served up some decent results across the wide range of tests put in front of it, with some good numbers in the CrystalDiskMark tests that match well against the claimed speeds from Samsung.

As for its top-line results, this Samsung SSD provided solid sequential speeds of 7208.62 MB/s for reads and 5762.93 MB/s for writes. The reads are in line with Samsung’s own claims, although the writes fall short of both its claims and those of rivals from Crucial and WD.

Advertisement
Samsung 990 Crucial P310 WD Black SN7100
CrystalDiskMark 8 Sequential Q8 Reads 7208.62 MB/s 7168.54 MB/s 7220.65 MB/s
CrystalDiskMark 8 Sequential Q8 Writes 5762.93 MB/s 6342.99 MB/s 6941.71 MB/s
CrystalDiskMark 8 Sequential Q1 Reads 3813.36 MB/s 3949.65 MB/s 4977.82 MB/s
CrystalDiskMark 8 Sequential Q1 Writes 5530.19 MB/s 4845.48 MB/s 6093.20 MB/s
CrystalDiskMark 8 Random 4K Q32 Reads 521.89 MB/s 622.54 MB/s 804.82 MB/s
CrystalDiskMark 8 Random 4K Q32 Writes 319.28 MB/s 342.03 MB/s 593.95 MB/s
CrystalDiskMark 8 Random 4K Q1 Reads 68.14 MB/s 50.19 MB/s 103.41 MB/s
CrystalDiskMark 8 Random 4K Q1 Writes 218.09 MB/s 221.11 MB/s 240.44 MB/s
FFXIV Endwalker Benchmark Loadtime 8.48 seconds 7.87 seconds 7.96 seconds
PCMark 10 QSD Benchmark 2217 2882 3043
PCMark 10 Data Drive Benchmark 3240 4190 4425
120GB Real World File Copy Test 40.3 seconds 37.3 seconds 46.66 seconds

However, its results at the Q1 level were somewhat behind in some areas, while its 4K results were simply average against the P310 and WD Black SN7100.

Advertisement

The 120GB file transfer took 40.3 seconds, which is in the ballpark against its rivals, thanks to an average transfer rate of 2.98GB/s. In addition, the Final Fantasy XIV Endwalker benchmark spat out a load time of 8.48 seconds. This is, again, an adequate result, although competing drives can be up to a second faster in my testing.

This Samsung drive falls down in the PCMark 10 Data Drive and Quick System Drive tests, with results more in line with other value drive efforts, such as the WD Blue SN580. It’s lower than a lot of other value PCIe 4.0 drives I’ve tested, which is a shame.

Advertisement

SQUIRREL_PLAYLIST_10208689

Should you buy it?

You want an adequate PCIe 4.0 SSD

There isn’t necessarily anything wrong with the 990, as its performance is decent for a more affordable PCIe 4.0 SSD.

Advertisement

You want a more affordable choice

Advertisement

The problem with this drive is its severely high price for the specs and performance on offer – it’s possible to get stronger results for less from a range of rivals.

Final Thoughts

The Samsung 990 is a solid PCIe 4.0 SSD with middle-of-the-pack sequential and random performance that’s decently brisk, alongside sensible capacity options and decent durability. For a more value-oriented drive, though, its retail pricing is some of the most baffling I’ve seen.

Advertisement

The WD Black SN7100 and the Crucial P310 outperform this drive in top-end speed, most random tasks and game load times and are more affordable than the 990, while Samsung’s own 990 Pro is faster and more feature-rich from resellers, in spite of being slightly older. It puts the 990 in a bit of a difficult position – there isn’t necessarily anything wrong with it, with cromulent performance for most folks, but its high price leaves a bit of a sting. For more options, check out our list of the best SSDs we’ve tested.

How We Test

Each SSD we test utilises a mix of both synthetic and real-world benchmark tests. On top of that, we also use a number of price-to-performance metrics, and monitor temperature and power-draw to determine the long-term stability and cost-effectiveness of the drive.

  • Each SSD is tested in a bespoke test PC across a number of different scenarios
  • SSD temperatures and power draw are monitored throughout the process

Advertisement

FAQs

Can I use the Samsung 990 in my PC?

Yes, as long as you’ve got an M.2 slot it will work in any PC.

Advertisement

Test Data

Full Specs

  Samsung 990 SSD Review
Manufacturer Samsung
Storage Capacity 1TB, 2TB
Size (Dimensions) 22 x 80 x 2.3 MM
Weight 9 G
Release Date 2026
First Reviewed Date 10/08/2026
Storage Type SSD
Read Speed 7250 MB/s
Write Speed 6450 MB/s
Interface PCIe 4.0 x4
Connector M.2
Heatset included? No
UK RRP TBC
USA RRP $269.99

Source link

Advertisement
Continue Reading

Tech

Private security firms will soon be allowed to hack overseas cybercriminals

Published

on

The Trump administration is recruiting private security firms to conduct federal government-authorized operations, including cyberattacks, against overseas-based criminal organizations that commit hacks on US persons, organizations, or government entities.

In a National Security Presidential Memorandum issued Thursday, US President Donald Trump directed the National Coordination Center (NCC), which operates under the Homeland Security Task Force, to develop a program for conducting specific cyber operations that combat foreign transnational criminal organizations (TCOs). The Departments of Justice and Homeland Security will provide oversight. The lynchpin of that program is bringing in private sector companies to participate.

Devil will be in the still-undefined details

A fact sheet that accompanied Thursday’s memo listed ransomware, sextortion schemes, phishing campaigns, financial fraud, and impersonation scams as activities eligible for private-sector security firms to target. The memo said such firms could “conduct Cyber Surveillance Operations and Cyber Effects Operations” against “cyber-enabled” TCOs. Such groups are defined as “any foreign group that conducts cyber-enabled crime against the United States Government, a United States person, or United States interests, and that is not an institutional part of a foreign government or wholly operated under a foreign government’s direction.”

The new program is the first time the federal government will authorize private companies to conduct offensive cyber operations against overseas hackers. The memo appears to permit companies participating in the program to use spyware or launch offensive attacks intended to destroy TCO data or systems. The memo doesn’t rule out certain types of offensive attacks, such as those that use encryption to lock targets out of their networks or performing distributed denial-of-service attacks. Up until now, the government has prohibited the private sector from taking such actions without court-authorized approval.

Advertisement

Source link

Continue Reading

Tech

Google drops Gemini 3.7 Flash model, and it’s ready to handle your chores with the Spark agent

Published

on

Google has launched Gemini 3.7 Flash, and one of its biggest upgrades is going straight to Gemini Spark.

The company says Spark is moving to the new model starting today, giving its cloud-based agent better tool use and stronger performance on multi-step tasks. Google specifically points to work such as consolidating files, drafting emails, and updating status documents across Google Workspace.

Spark could get a lot better at doing things for you

Gemini Spark can already continue working after a laptop is closed or a phone is locked, and it can operate across services such as Gmail, Drive, Docs, Calendar, Keep, and Tasks.

Google recently pushed those abilities further through Chrome. Spark can use accounts you are already signed into and passwords saved in the browser to navigate websites, compare flight options, schedule apartment viewings, and begin bookings. It still hands control back before sensitive actions such as payments.

Advertisement

Gemini 3.7 Flash is supposed to make those workflows more reliable. Google says the model spends more effort on planning and tool calls, adapts better when it runs into roadblocks, and should need fewer retries and less manual intervention.

The benchmark results point in the same direction. Gemini 3.7 Flash scored 30.4% on AutomationBench, up from 17% for Gemini 3.6 Flash. It also came in ahead of GPT-5.6 Terra at 23.6% and Claude Sonnet 5 at 10.7%.

Gemini 3.7 Flash is stronger beyond Spark too

Google also posted a sizable jump on document-heavy work. Gemini 3.7 Flash scored 34% on GDP.pdf, compared to 22% for Gemini 3.6 Flash, 28% for Claude Sonnet 5, and 24.7% for GPT-5.6 Terra.

The pricing is what makes those gains more interesting. Gemini 3.7 Flash costs $0.75 per million input tokens and $3.75 per million output tokens under introductory pricing through December 31. It is significantly cheaper than Claude Sonnet 5 at $2 and $10, and GPT-5.6 Terra at $2 and $12.

The model also posted improvements on coding-focused benchmarks. Gemini 3.7 Flash scored 43.6% on FrontierCode, edging out Claude Sonnet 5 at 42.7% and GPT-5.6 Terra at 41.3%, while also improving over Gemini 3.6 Flash on DeepSWE.

For Spark, Gemini 3.7 Flash looks like a meaningful upgrade, especially if the improvements in tool use and multi-step planning translate into fewer failed tasks and less manual intervention. Power users, however, are still waiting for Google’s more ambitious Gemini 3.5 Pro model, which was previewed at Google I/O in May but has yet to ship after missing its expected June rollout.

Advertisement

Source link

Continue Reading

Tech

The Safety Reckoning Inside OpenAI

Published

on

OpenAI’s leaders are rallying workers to respond to one of the largest crises in the company’s history—which spans across its AI safety, cybersecurity, and alignment divisions. The ChatGPT-maker says it has slowed down research, spent millions of dollars, and told several teams to drop everything to focus on investigating a set of rogue AI agents that breached the platform Hugging Face in a quest to complete an internal security test.

OpenAI is expected to release a comprehensive postmortem detailing the incident in the coming days. However, the Hugging Face incident has inspired OpenAI leaders and employees to examine how the AI lab’s culture may have enabled this incident in the first place.

Multiple current and former OpenAI employees, who spoke on the condition of anonymity to discuss private internal matters, tell WIRED they believe competitive pressures to quickly ship new AI models and products have made it difficult for staffers to sufficiently prioritize safety, security, and alignment.

“We’re reaching new levels of model capability that require more robust training, alignment, safety and security testing, deployment practices, and governance—as demonstrated by the work we’re doing to prepare Astra and future models,” said OpenAI president and cofounder Greg Brockman in a statement to WIRED. “We feel the weight of deploying our models and products responsibly, and a lot of that starts with the changes we’ve made to more deeply integrate research, safety, and security into frontier-model development from the start.”

Advertisement

This is far from the first time OpenAI employees have raised such concerns. Back in 2024, OpenAI’s then head of alignment Jan Leike left to join Anthropic, warning on his way that safety was taking a back seat to shiny products. Two years later, the Hugging Face attack represents a watershed moment for the AI industry, demonstrating that AI agents today can cause real-world harm when safety, security, and alignment aren’t properly accounted for.

“We are responding to this with the utmost severity,” said Michael Dalton, an OpenAI security and infrastructure engineer, during a talk at the Black Hat cybersecurity conference last week. “What I would internalize is that AI-orchestrated, fully automated offensive attacks are real now. The actions we have discussed today were an unintended side effect of running evaluations on frontier AI.”

Some OpenAI employees told WIRED they are optimistic this incident will inspire genuine change within the company. OpenAI has committed to slowing the release of future AI models and has been especially forthcoming about areas where its mitigations fell short. Boaz Barak, a researcher who coleads OpenAI’s safety advisory group, said in a post on X that addressing the situation “requires not just fixing some issues but also changing our culture.”

In their Black Hat talk, OpenAI security engineers Dalton and Eric Wallace said that the Hugging Face incident started in May when, unbeknownst to the company, several AI agents thought to be operating within isolated testing environments gained access to the internet and convened on a covert message board to coordinate with one another.

Advertisement

OpenAI would not discover the message board until July, when it learned that the AI agents had hacked into multiple services to try to achieve their larger goal of breaching Hugging Face’s platform, which they believed may contain answers to the security tests they were trying to solve.

“They were incredibly sloppy. If you’re serious about this, your AI shouldn’t be able to break out onto the internet and then do it again right afterward,” says one former OpenAI employee who requested anonymity to speak with WIRED. “This was the biggest safety incident in OpenAI’s history.”

The New Guard

Weeks before OpenAI discovered the Hugging Face incident, WIRED reported that the company had begun a reorganization to combine its safety and core research teams, which led to the departure of its then safety leader Johannes Heidecke.

Sandhini Agarwal, who led AI safety teams at OpenAI, also left the company in July after more than six years, according to her LinkedIn. Agarwal did not immediately respond to WIRED’s request for comment.

Advertisement

Source link

Continue Reading

Trending

Copyright © 2025