Connect with us

Tech

IBM, OpenAI join forces to scale AI adoption and boost security

Published

on

Specialised engineers and consultants trained through OpenAI’s partner network will work directly with clients on solutions.

OpenAI – in its latest effort to garner enterprise support – has teamed up with IBM to provide organisations with its tools to scale AI deployment.

The deal will see the two companies create industry-specific, go-to-market initiatives targeted at financial services, government, telecommunications and retail sectors, and comes at a time when AI investments are surging multi-fold despite only a fraction of businesses seeing real returns.

“While enterprises are rapidly investing in AI, they are looking for practical ways to apply it across their core operations to deliver measurable business outcomes and create new commercial models,” said Andy Baldwin, the global senior vice-president at IBM Consulting.

Advertisement

“The challenge is not access to AI technologies – it’s integrating AI securely and at scale into complex enterprise environments and workflows.”

The partnership embeds OpenAI’s latest frontier models, including GPT-5.6 and products like Codex and ChatGPT Work, into IBM’s AI platform.

IBM said it will deploy units of specialised engineers and consultants trained through OpenAI’s partner network to work directly with clients to accelerate AI implementation across complex business workflows and highly regulated environments.

The technology giant is also launching a dedicated channel through which more of its consultants and engineers can be certified under the partner network.

Advertisement

The two also said they will expand their collaboration as part of OpenAI’s cybersecurity partnership programme Daybreak by combining the AI giant’s frontier capabilities with IBM’s multi-agent-powered service for delivering decision-making and intelligence.

With this, the collaborators aim to help clients manage both cyber and AI model risk, including application-layer vulnerabilities, governance gaps and operational risks that limits AI adoption, they said.

“The organisations pulling ahead with AI are the ones turning it into a trusted part of how their business operates,” said Denise Dresser, the chief revenue officer at OpenAI.

Earlier this year, IBM set aside more than $10bn for its lofty quantum plans that it hopes will deliver the world’s first large-scale, fault-tolerant quantum computer by 2029.

Advertisement

Don’t miss out on the knowledge you need to succeed. Sign up for the Daily Brief, Silicon Republic’s digest of need-to-know sci-tech news.

Source link

Advertisement
Continue Reading
Click to comment

You must be logged in to post a comment Login

Leave a Reply

Tech

Say goodbye to Chronicle. ChatGPT’s new Computer History feature does it better

Published

on

OpenAI just rolled out Computer History for the ChatGPT desktop app, and if you found Chronicle interesting but a little too screenshot-happy, this update should ease your mind.

What does Computer History do?

Instead of taking screenshots as Chronicle did, Computer History tracks your clicks, keystrokes, and app switches through macOS’s accessibility system, then turns all that activity into text summaries and a timeline you can scroll through. Think of it as ChatGPT quietly taking notes on your day so it can help you out later.

ChatGPT can now remember your activity across the apps and websites on your computer.

With Computer History in the desktop app, future interactions feel more personalized and require less explanation. pic.twitter.com/WHZPxPp31R

— OpenAI (@OpenAI) August 13, 2026

Advertisement

Once it’s turned on, you can ask ChatGPT things like what you were working on before your last break, or where that proposal document went. It digs through your recent activity and points you in the right direction instead of making you reconstruct your entire morning from memory.

If you find yourself doing the same task over and over, Computer History can notice the pattern and suggest turning it into a skill or automation. You just review the suggestion and let Codex build it out for you.

How do you turn it on, and is it actually private?

Computer History is off by default, and it stays that way until you flip the switch yourself. To turn it on, head to Settings → Integrations, and select Computer History to get started. You’ll also need Memories turned on, since that’s what lets ChatGPT actually use this context across your chats.

You get to pick exactly which apps and websites are allowed to contribute, and private browsing is never touched. There’s also a menu bar shortcut to pause collection whenever you want a break from being tracked, along with the option to delete your history at any time.

Advertisement

If you’re on a Business or Enterprise workspace, your admin needs to grant access first, and even then, you still have to opt in yourself. Pro users can enable it from the settings.

Overall, this feels like a solid privacy-first upgrade over Chronicle, and I like that OpenAI is giving users this much control over what gets remembered.

Source link

Advertisement
Continue Reading

Tech

Akira hackers disable EDR with Safe Mode, steal data but fail to encrypt

Published

on

Akira hackers disable EDR with Safe Mode, steal data but fail to encrypt

An Akira ransomware affiliate disabled the endpoint detection and response (EDR) solution on a compromised system by restarting the machine into Safe Mode with Networking.

The attack occurred on August 4 after the hacker obtained initial access through an exposed SonicWall VPN device without multi-factor authentication (MFA).

Managed detection and response (MDR) services company Huntress says that roughly two hours after a successful VPN login, the attacker connected to the domain controller via RDP, enumerated Active Directory users and computers, and then moved to an application server.

image

They used WinRAR to archive mapped file shares and the s5cmd command-line tool to upload the stolen data to an attacker-controlled S3 bucket, before installing AnyDesk for remote access.

At that stage, the attacker used AnyDesk to force the compromised host to boot into Safe Mode with Networking and disable both the Huntress agent and Microsoft Defender’s real-time protection.

Advertisement

Safe Mode is a Windows startup state designed for troubleshooting and diagnostic operations. It starts Windows with a limited set of drivers and services, generally preventing most third-party software and services from loading.

For 10 minutes while in Safe Mode, “the host had no working EDR, and AV was blinded,” Huntress says.

Meanwhile, the attackers added AnyDesk to Windows’ Safe Mode registry, allowing it to start after reboot and retain their remote access to the breached machine.

However, when they attempted to launch the main ransomware payload (akira.exe) via AnyDesk in Safe Mode, it failed to execute as the system reported low virtual memory and generated out-of-memory and PowerShell errors.

Advertisement
Akira attack flow
Akira ransomware attack flow
Source: Huntress

A scheduled Defender scan eventually detected the Akira executable, even if real-time protection was disabled in Safe Mode, but the security tool could not remove it while the machine remained in that mode.

Defender quarantined the file only after the attacker rebooted the system into normal mode, which restored real-time protection.

Despite the failure to encrypt files, the Akira operator still managed to steal credentials and files for data extortion, all in less than five hours from initial access.

Huntress notes that other ransomware families, such as Snatch and AvosLocker, have used this tactic for years, but this incident marks the first time the company observed it in an Akira attack.

The researchers recommend adding MFA to all VPN accounts, placing credential-spraying detection measures, and monitoring for Safe Mode boot configuration changes or remote-access tools being added to the Safe Mode service registry.

Advertisement

article image

Overall prevention scores can hide what happens after initial access. Once attackers are using valid credentials, prevention drops sharply.

The Blue Report 2026 measures defenses technique by technique across 338 million simulations run in customer production environments.

Get the report

Source link

Continue Reading

Tech

Investors eye $2trn Anthropic valuation in record IPO

Published

on

Fervour for a massive valuation comes as a result of Anthropic’s rapidly growing revenue, sources told the FT.

Investors expect Anthropic to be valued at $2trn or more in its October initial public offering, potentially doubling initial targets of $1trn and dwarfing SpaceX as the largest public listing in history.

Several of the company’s investors told the Financial Times (FT) that the fervour for a massive valuation comes as a result of Anthropic’s rapidly growing revenue, which reportedly hit nearly $11bn in the second quarter of this year – more than double the $4.8bn of the first quarter.

The AI giant, which filed to go public in June, is yet to fix on a valuation, the FT reported.

Advertisement

The company’s backers expect the five-year-old AI giant to reach an annualised revenue of between $100bn and $120bn this year.

Anthropic’s Claude AI products are repeat headline-makers as it competes for leadership in the space with its biggest rival OpenAI – which also hopes to go public – and the more recent crop of Chinese-made models taking the industry by storm with their cheaper alternatives.

Confident backers – including venture capitalists, other industry giants and institutional investors – have poured nearly $100bn into Anthropic just this year, fuelling the business as it looks to build its own AI chips to keep up with surging demand.

Last month, AMD pledged $5bn to Anthropic in a deal that also allows the AI giant access to 2GW of AMD’s latest-generation chips. In April, Amazon announced plans to pour up to $25bn into Anthropic, which in turn pledged to spend around $100bn over the next 10 years on the e-commerce juggernaut’s cloud technologies.

Advertisement

While Anthropic does not share details of how many use Claude, Statista placed it at around 245m monthly users as of June this year. Comparatively, OpenAI’s ChatGPT reached 1bn users in May.

Despite this, Anthropic trumped OpenAI’s valuation earlier this year, owing to its growing share of the more lucrative enterprise sector, where it has been capturing a higher volume of first-time users.

Earlier this year, the US government temporarily banned Anthropic from exporting two of its highly capable cybersecurity models over security concerns.

While that ban was eventually lifted after a little more than two weeks, sources told the FT that Anthropic’s June revenue suffered as a result. The company, however, rebounded at an “extraordinary rate” once the ban was lifted, they said.

Advertisement

Anthropic is separately embroiled in an ongoing legal battle with the US government over the banning of its products for official use.

Don’t miss out on the knowledge you need to succeed. Sign up for the Daily Brief, Silicon Republic’s digest of need-to-know sci-tech news.

Source link

Advertisement
Continue Reading

Tech

Apple sends new ‘Threat Notification’ alerts over mercenary spyware attacks

Published

on

Apple

You’re not alone if you just received an “Apple Threat Notification” saying it detected a “mercenary spyware attack targeted at your iPhone.”

Some users on Reddit are reporting that they received these alerts today after Apple sent out a new batch of threat notifications on August 13, but the feature itself is not new.

Apple Threat
Apple sent a new batch of alerts to users on August 13

Source: Reddit

Apple has been sending these threat notifications multiple times a year since 2021, when it detects highly targeted mercenary spyware attacks.

image

It’s also worth pointing out that Apple does not identify the spyware behind individual alerts, so there’s no evidence that today’s notifications are specifically related to Pegasus.

However, Apple itself cites NSO Group’s Pegasus as an example of mercenary spyware historically associated with this type of attack, and forensic investigations into previous Apple threat notifications have confirmed Pegasus infections in some cases.

In a support document, Apple previously confirmed it sends threat notifications to users in more than 150 countries after detecting highly targeted mercenary spyware attacks against specific iPhone users.

Advertisement

The list of potential targets includes journalists, activists, politicians, and diplomats, who have historically been among those targeted by this type of spyware.

These attacks are expensive, highly sophisticated, and typically aimed at a very small number of people.

“Mercenary spyware attacks cost millions of dollars and often have a short shelf life, making them much harder to detect and prevent,” Apple explained.

“The vast majority of users will never be targeted by such attacks.”

Advertisement

The company does not attribute individual alerts to a specific government, company, or geographical region.

Apple says threat notifications should be taken seriously

Apple relies on its own threat intelligence and investigations to identify suspected mercenary spyware activity, which means these notifications are “high-confidence alerts” and not just a regular warning.

“Although our investigations can never achieve absolute certainty, Apple threat notifications are high-confidence alerts that a user has been individually targeted by a mercenary spyware attack, and should be taken very seriously,” Apple noted.

“We are unable to provide information about what causes us to issue threat notifications, as that may help mercenary spyware attackers adapt their behavior to evade detection in the future.”

Advertisement

If Apple detects this activity, it sends an email and iMessage notification to the email addresses and phone numbers associated with the user’s Apple Account. 

The emails are usually from threat-notifications@email.apple.com, and Apple also warns users about fake versions of these alerts.

You can verify whether a threat notification is genuine because Apple will not ask you to click a link, open a file, install an app or profile, or provide an Apple Account password or verification code.

You can also check the alert by signing in directly to account.apple.com. If Apple sent you a threat notification, it will appear at the top of the page after you’re logged in.

Advertisement

If you believe you’ve been affected, you should enable Lockdown Mode and reach out to a cybersecurity expert.

Apple recommends taking these alerts seriously because receiving one means it has high confidence that the user was individually targeted.

BleepingComputer has contacted Apple for a statement on the Threat Notifications, but we have not received a response at the time of publication.


article image

Overall prevention scores can hide what happens after initial access. Once attackers are using valid credentials, prevention drops sharply.

The Blue Report 2026 measures defenses technique by technique across 338 million simulations run in customer production environments.

Advertisement

Get the report

Source link

Continue Reading

Tech

Thinking of a career in fintech? Then brush up on these skills

Published

on

The fintech space is moving along at an alarming pace, and these skills for staying informed and capable should be kept in mind.

In the rapidly evolving STEM space, it can sometimes feel as though the moment you have mastered a new skill or overcome an obstacle, a larger, more pressing challenge takes its place. It can be extraordinarily difficult to stay up to date on industry changes, but for ambitious jobseekers and professionals, falling behind is not an option.

The fintech space is no different. Modernisation, particularly in the area of digital assets and financial infrastructure, has resulted in an ecosystem where the vast majority of the work has moved online and the skills needed to stay informed have become increasingly complex.  

But that doesn’t mean that it is impossible to do so – in fact, now more than ever there are a myriad of fun, innovative and convenient ways to upskill, without breaking the bank. 

Advertisement

With that in mind, as part of SiliconRepublic.com’s August fintech and financial infrastructure coverage, here are some of the skills that you should definitely keep on your radar if you envision a long and successful career in this space. 

Keep it real

To start off, it is critical that in an evolved, automated and complex environment, where often you are dealing with deeply sensitive information, you have an arsenal of interpersonal skills. Many professionals or jobseekers in finance might assume that to get ahead, all you need are the technical skills and a decent understanding of the lingo. 

While this is an equally important component of a fintech or finance-based career, professionals looking for a robust skillset should ensure that they have strong communication skills, that they can address challenges and concerns with empathy, that they engage with private information ethically and that they can adapt to a rapidly changing sector. 

This adaptability will be evident in how you engage with new policies, tools, techniques and shifts in workplace expectations. Remember, your education, experience and qualifications are often what gets you the interview, but it is the soft skills and your attitude that indicate whether you are a good fit, not just for the industry but also the company or institution you are applying for, which will have its own set of values. 

Advertisement

Keep it relevant

Financial literacy is the bedrock of an education or career in fintech and financial infrastructure, as it enables the student or professional to understand core concepts and make informed decisions. Financial literacy gives those who possess it an evergreen, foundational knowledge of key topics such as budgeting, investing and borrowing, and it allows individuals to better understand an evolving marketplace. 

Online courses teaching financial literacy are an ideal resource, as are books, videos and essays covering the subject. Fintech applications and community organisations can also give professionals or students a learning outlet that feels more ‘real’, in comparison to a book or course, as activity occurs in-person or in real-time. While the industry is always changing, many of the core concepts stay the same, so don’t get distracted by the shiny new tools to the point that you forget the basics. 

Keep it skilful 

Since the evolution of coding towards a landscape where ‘vibe-coding’ and AI-powered programming do much of the work, some might consider traditional coding skills less important now. However, this couldn’t be further from the truth. As more and more roles become cross-collaborative and skill expectations evolve, coding capabilities give finance professionals an edge in an ecosystem that thrives on niche or under-represented skills.

Languages such as Python, Java and JavaScript are often used in the development of fintech platforms. Having the know-how to build an app and then deploy, maintain and use it has the potential to create an invaluable employee who is useful across multiple areas of an organisation. And it looks great on a CV or in a job interview, as it shows you are not confined by the typical skills expectations of a role. 

Advertisement

Keep it standing 

What is the point in having people skills, foundational skills or coding skills geared towards a fintech or finance-based career if your knowledge of the sector’s infrastructure is inherently weak?

Students and professionals eager to develop a strong career in this space should ensure that they have a strong understanding of the industry’s infrastructure. That includes, but isn’t limited to, banking systems (traditional and neo-banks), payment platforms, and the risk and compliance organisations that govern it all. Areas to prioritise include financial analysis, finance modelling skills, increased awareness of regulatory and political factors, and complex project management, among others.  

All in all, as long as you can say you are continuously learning and can show that you aren’t allowing your skills to atrophy, a good employer, who is serious about ensuring a candidate is a good fit, will reconsider the potential you bring. 

Don’t miss out on the knowledge you need to succeed. Sign up for the Daily Brief, Silicon Republic’s digest of need-to-know sci-tech news.

Advertisement

Source link

Continue Reading

Tech

Apple CEO Tim Cook doesn’t want to define his legacy

Published

on

Tim Cook will step down as Apple CEO on September 1, and he has shared that others will define his legacy, but he’d like to be viewed as a “good and decent man.”

John Ternus will be Apple’s CEO starting on September 1 and Tim Cook will become the Executive Chairman. Cook’s legacy will be defined and litigated for decades to come, but he won’t be entering the conversation himself.

According to an interview on CNBC, Cook doesn’t want to try to define his own legacy. He says that others will define his legacy for him, but that he hopes that he will be seen as a “good and decent man.”

The clip is short, but here it is:

Advertisement

Cook will still be around to handle matters as a special government liaison for Apple. While he won’t be appearing in keynotes, he will likely be seen at select press events for some time to come.

While there will always be some debate around Cook and his time as CEO, but his record is clear. As was shared with his last earnings call, the company earned $109 billion for the quarter, which surpassed what the company earned in the entire first year Cook was CEO.

Source link

Advertisement
Continue Reading

Tech

NYT Strands hints and answers for Friday, August 14 (game #894)

Published

on

Looking for a different day?

A new NYT Strands puzzle appears at midnight each day for your time zone – which means that some people are always playing ‘today’s game’ while others are playing ‘yesterday’s’. If you’re looking for Thursday’s puzzle instead then click here: NYT Strands hints and answers for Thursday, August 13 (game #893).

Strands is the NYT’s latest word game after the likes of Wordle, Spelling Bee and Connections – and it’s great fun. It can be difficult, though, so read on for my Strands hints.

Want more word-based fun? Then check out my NYT Connections today and Quordle today pages for hints and answers for those games, and Marc’s Wordle today page for the original viral word game.

Advertisement

Source link

Advertisement
Continue Reading

Tech

Darksiders 4 and the next Kingdom Come game could arrive by March 2028

Published

on

Fans waiting for Darksiders 4 and the next Kingdom Come game have finally been given something resembling a release window. Unfortunately, it is a rather large one. As part of its Q1 FY 2026/27 report, Embracer has placed both games in its 2027/28 fiscal-year pipeline, meaning they are currently expected to launch sometime between April 2027 and March 2028.

Darksiders 4 is still a ways away

Darksiders 4 was announced in 2025, but Embracer has remained fairly quiet about the project since then. The company’s latest plans confirm that it is part of the major releases expected during FY 2027/28, alongside titles including Tomb Raider: Catalyst, which should follow up on Tomb Raider: Legacy of Atlantis set to launch in 2027.

That doesn’t mean Darksiders 4 has a March 2028 release date. Fiscal 2027/28 simply runs from April 2027 through March 2028, so the actual launch could happen at any point within that period. Embracer’s previous release documents also list Darksiders 4 for PC, PS5 and Xbox Series X|S.

Advertisement

The other interesting part is the next Kingdom Come game. Embracer’s annual report says Warhorse Studios is working on a new release in the franchise, while the company’s current pipeline places it in the same FY 2027/28 window. Importantly, Embracer has not officially called it Kingdom Come: Deliverance III, so it is better to think of this simply as the next game in the franchise for now. Warhorse is also developing a separate Middle-earth project, meaning the studio has multiple major projects in the works.

That’s a lot of games for 2027-28

The timing makes sense within Embracer’s broader strategy. The company says its Fellowship Entertainment division, which houses major franchises including Kingdom Come, Darksiders, Tomb Raider and Dead Island, is targeting a higher cadence of major releases from FY 2027/28 onward. It expects at least two major games per year with full economics starting in that period.

So, yes, Darksiders 4 and the next Kingdom Come game are coming. Just don’t clear the calendar yet. With the current window stretching all the way from April 2027 to March 2028, fans could still be waiting quite a while before either game actually lands

Source link

Advertisement
Continue Reading

Tech

Writer says its new Palmyra X6 model cuts AI agent costs by 52% as token spending surges

Published

on

Writer, the enterprise AI agent platform used by Fortune 500 companies including Accenture, Uber, and Vanguard, released its new flagship model Palmyra X6 today, alongside a rebuilt agent orchestration “harness” and new governance tools designed to give IT leaders control over runaway token spending.

The headline numbers are striking: Writer says its agent product now operates at an average 52% lower cost, with a 48% improvement in speed and a 10% improvement in quality when paired with Palmyra X6. But the more consequential story may be how the company got there — and what its choices reveal about where the enterprise AI market is heading.

Palmyra X6 is not trained from scratch. It is a post-trained version of GLM-5.2, the open-weight mixture-of-experts model from Beijing-based Z.ai, formerly Zhipu AI — a fact Writer discloses openly in its technical report, and one that places the San Francisco company at the center of one of the industry’s most charged debates: whether American enterprises should build on Chinese open-source foundations.

“This model is in no way, shape, or form connected to any of its original developers. It is fully run on our U.S. infrastructure,” Matan-Paul Shetrit, Writer’s director of product management, told VentureBeat in an exclusive interview ahead of the announcement.

Advertisement

Dan Bikel, who leads Writer’s AI research, put it more bluntly: “It’s very much a Palmyra model, and we just happen to grab the floating point numbers as the starting point, and train from there.”

Why AI agents are blowing up enterprise budgets in ways chatbots never did

Writer’s announcement lands at a moment when the economics of agentic AI have moved to the center of enterprise buying decisions. Unlike a chatbot, which typically generates one answer per user request, an AI agent turns a single request into repeated rounds of planning, retrieval, tool calls, validation, and retries — with every loop consuming metered tokens. The user sees one answer; the invoice reflects the entire loop.

The scale of the problem is becoming clear. Goldman Sachs forecasts that token consumption will multiply 24 times between 2026 and 2030, reaching 120 quadrillion tokens per month, driven not by more people asking questions but by always-on enterprise agents. The same analysis warned that falling per-token prices do not guarantee falling bills: if an agentic task draws 20 times more tokens while unit prices fall 75%, total charges still rise fivefold.

“The enterprise wants token consumption to explode — it means adoption is happening — but they need costs to flatten,” said Waseem AlShikh, Writer’s CTO and co-founder, in a statement.

Advertisement

Shetrit framed the cost problem as the primary obstacle to enterprise AI adoption — more so than model capability itself. “The biggest barrier today to enterprise expansion using AI is actually not model capabilities in most cases; it’s actually the cost around them,” he said. “The reality today is, in most cases, the alternative for AI is not another AI, it is human labor.”

Asked whether cutting customers’ token consumption would cannibalize Writer’s own per-token revenue, Shetrit rejected the premise. “Reducing the cost is not hurting my bottom line. It’s actually expanding it, because it’s expanding the TAM of opportunity within an organization,” he said, arguing that lower per-task costs unlock workflows enterprises would otherwise never automate. That argument echoes a pattern familiar from the cloud era, where unit prices fell for a decade while total bills rose as consumption expanded — a dynamic Writer is explicitly betting will repeat with agents, and betting it can profit from.

Inside Palmyra X6: how 626 training examples fine-tuned a 744-billion-parameter model

Palmyra X6 is a 744-billion-parameter mixture-of-experts model with roughly 40 billion active parameters per token, inheriting GLM-5.2’s architecture unchanged, according to Writer’s technical report. The company’s contribution is a deliberately conservative post-training recipe: a technique called anchored supervised fine-tuning (ASFT), applied to a remarkably small corpus of just 626 curated synthetic agentic trajectories, trained for a single epoch at a low learning rate.

The tiny dataset is the point, not a limitation. ASFT pairs a token-weighting scheme with a KL-divergence “anchor” that penalizes the fine-tuned model for drifting too far from a frozen copy of the base model — teaching new tool-use behaviors without eroding the general capabilities the base already has. Writer also swapped the standard Adam optimizer for Muon, a newer method that treats weight matrices as geometric objects, on the model’s core weight matrices.

Advertisement

“There’s a whole string of papers following a quote-unquote ‘less is more’” philosophy, Bikel said, referencing research showing that “small, extremely high quality data sets go a really long way.” He added: “That’s the philosophy — one of the philosophies — that we followed when building this model, and it showed. It allowed us to optimize for our customers at lower cost to do the work of optimization, and that ultimately yielded a lower cost model for us and for them.”

The training data itself is fully synthetic — every plan, tool call, and final answer machine-generated by teacher models, then filtered through structural quality gates, a model-based verifier, and a two-model LLM judging panel before entering training. That continues a long-standing Writer practice: the company’s Palmyra X 004 was trained almost entirely on synthetic data for roughly $700,000 back in 2024, as TechCrunch reporte at the time, and Palmyra X5 required about $1 million in GPU hours, according to SiliconANGLE.

On Writer’s internal evaluations — nine capabilities spanning grounding and retrieval, tool use, content generation, sub-agent delegation, and brand voice — X6 scored an average of 0.87 out of 1.00, edging out Anthropic’s Claude Opus 4.8 (0.86), Claude Sonnet 4.6 (0.85), OpenAI’s GPT-5.5 (0.80), and Google’s Gemini 3.1 (0.77). The price gap is the real differentiator: Writer prices X6 at $2 per million input tokens and $8 per million output tokens, versus 15/75 for Opus 4.8. The company says X6 completes tasks in 26 seconds on average and can work unattended toward a single goal for up to eight hours.

Writer is candid that internal benchmarks invite skepticism. Asked directly whether the company would publish its methodology after grading its own homework, Bikel said the technical report covers “both the protocol we used to do our public benchmarking as well as our internal evaluations.” He described public benchmarks as sanity checks rather than targets: “We do things like public benchmarks to let us know that we’re climbing the right hill and that we don’t have any sort of huge gaps, but we don’t slavishly follow them either, because that’s not really serving our customers.”

Advertisement

The China question: what building on GLM-5.2 means for enterprise security and trust

Writer’s choice of base model would have been unthinkable for an American enterprise vendor two years ago. Today it reflects a market reality: GLM-5.2, released in June under the permissive MIT license, is arguably the most capable openly available model in the world. Independent analysis house Artificial Analysis scored it at 51 on its Intelligence Index — ahead of DeepSeek V4 Pro, Kimi K2.6, and even some of Google’s Gemini models on agentic tasks — while undercutting U.S. flagship API pricing many times over, as European tech outlet Trending Topics reported. Writer’s press release calls it “the strongest available open-weight model.”

The open-weight surge carries genuine baggage. An August report from AI safety nonprofit SaferAI found that GLM-5.2 refused none of the offensive cyber or biology tasks it was given via Z.ai’s public API, and that Z.ai published no safety framework or pre-deployment risk assessment — a gap that widens once anyone can download and modify the weights.

Writer’s answer is that provenance and post-training matter more than origin. Bikel emphasized that the company “grabbed the weights off of the U.S. Hugging Face” and trained entirely on American infrastructure; the technical report states all datasets were synthesized and stored in the U.S., and all training hardware was located in the U.S.

The company also ran what it describes as an unusually rigorous, pre-registered model-risk evaluation covering political bias, censorship, factuality, and refusal behavior — 19,674 evaluated responses scored by blinded judges — comparing X6 against its GLM-5.2 base and four frontier control models.

Advertisement

On the Washington Post’s ModelSlant political-bias evaluation, Writer says X6 presented both sides of hot-button questions 80% of the time, the highest rate of any model tested, and answered politically sensitive prompts that DeepSeek V4 refused outright. On the FORTRESS adversarial safety benchmark, X6 with its deployment system message scored 8.6 points higher on adversarial safety than the raw GLM-5.2 base, at negligible cost to benign helpfulness.

“We’ve run extensive benchmarking around bias, around censorship,” Shetrit said, “and the work Dan and the team has done has actually proven that this model is actually significantly better than not just open source alternatives, but any closed source alternative in the market at the time of the benchmarking.” 

The report does hedge in one notable place: while English-language behavior showed no statistically robust political asymmetry, “the behavior was shown to vary by language” — a candid admission that 626 fine-tuning trajectories do not scrub every trace of a base model’s training.

The harness effect: why orchestration may matter more than the model itself

Perhaps the most strategically interesting claim in Writer’s announcement has nothing to do with Palmyra X6 at all. The company says its rebuilt Writer Agent harness — the orchestration layer that plans tasks, batches work, delegates to sub-agents, and manages context — cuts costs by 41% and completes tasks 44% faster across every model it tested, including third-party models from Anthropic and OpenAI, while maintaining quality. Writer published the finding in an accompanying research paper on what it calls “The Harness Effect.”

Advertisement

That raises an obvious question, which VentureBeat put to the company: if the harness alone delivers most of the savings on any model, why build a model at all?

Shetrit’s answer was about control. “I cannot control if a lab deprecates their model. I cannot control what data they use in their model,” he said. “Where when I build the model, I have significant moral control, and I can answer the tough questions that enterprise customers ask me.”

Bikel added that the model and harness were developed together: “This model was built and essentially co-evolved with the harness… We know that we have a flagship product, Writer Agent. We want that to work really, really well with this model, and sure enough, it does. And we take that into account during model development, and that’s something that is not possible if you don’t build your own model.”

Notably, Writer is simultaneously hedging. With this release, the company extends multi-model support to Writer Agent, letting admins enable models from Anthropic, OpenAI, and cloud providers including Microsoft Azure, AWS Bedrock, and Nvidia NIM — even image-generation models, a category Writer does not build. The message to CIOs is disarmingly simple: use our model because it is cheapest and best for your workflows, but the platform saves you money either way.

Advertisement

New governance tools aim to end surprise AI bills before they start

The third leg of the release targets a quieter enterprise pain point: nobody in the C-suite knows what the agents are spending. New governance tools give administrators a centralized view of agent usage across the business, per-workflow analytics for the company’s shareable “Playbooks” and “Skills” automations, and consumption controls with alerts and spending limits.

Asked whether the introduction of spending controls implied that customers had been receiving surprise bills, Shetrit reframed it as an adoption enabler rather than damage control. “How do we build the tools to allow you as the CIO, CISO in a company, to feel comfortable both on the security and spend, so you can expand AI usage in your organization,” he said. In his telling, visibility is what lets leaders say yes: businesses with clear cost data “are actually looking to expand AI adoption to use cases that they would never have touched before.”

The feature set tracks a broader shift in how enterprises budget for AI. As Forbes analysis of the token price wars argued, sophisticated buyers are learning to model cost per successful task — counting retries, tool calls, and escalations — rather than multiplying expected calls by the advertised rate card. Writer is effectively productizing that discipline, turning what has been a finance-team spreadsheet exercise into a native platform capability.

It also completes a governance arc the company has been building for over a year. Writer shipped its unified agent experience with admin controls last November, then added agent Skills and workflow analytics in March, according to earlier company announcements. Thursday’s release closes the loop by attaching a price tag — and a spending limit — to every workflow.

Advertisement

Writer, founded in 2020 by May Habib and Waseem AlShikh, raised $200 million at a $1.9 billion valuation in late 2024, and has built its business on regulated, high-stakes deployments rather than consumer scale. Shetrit made no apology for the narrowness of that focus. “The privilege of working and focusing on enterprise use cases is that I don’t need my model to be able to write a French sonnet,” he said. “When you don’t try to do everything, you can focus on your customer problem and needs.”

He was equally direct about identity: “We are not a research lab converted to a consumer product now dabbling in enterprise. We are first and foremost an enterprise company that serves enterprise customers, and we evaluate our decisions within that lens. Which means, if we think building things from scratch is the right decision, that’s what we will do. But if we think there are other alternatives out there in the market that serve our customers better, that’s what we will do.”

That pragmatism may be the release’s most important signal. A well-capitalized American AI company with five years of model-building experience has concluded that the frontier of value no longer lies in pretraining, but in the last mile: post-training open weights, engineering the harness around them, and handing the CFO a dashboard. If Writer is right, the frontier labs’ moat narrows to the workloads where quality genuinely justifies a sevenfold price premium — and for everything else, the winning model is the one somebody else paid to pretrain.

In an industry that has spent three years arguing about whose model is smartest, Writer is making a different wager: the enterprise AI race won’t be won by the company with the best floating point numbers, but by the one that knows what to do with them.

Advertisement

Source link

Continue Reading

Tech

Google’s cheap model is now two versions ahead of its flagship

Published

on

Google has released Gemini 3.7 Flash with sharp gains on coding benchmarks and introductory pricing of $0.75 per million input tokens. Gemini 3.5 Pro remains months behind schedule, and Google will not say whether it is still coming.

Google has released Gemini 3.7 Flash, and still will not say when its flagship model is coming. The workhorse model posts large gains on coding, which is the ground Google has been losing. Gemini 3.5 Pro, the model this one is meant to sit beneath, remains months behind schedule.

The numbers are real. On FrontierCode 1.1 the new model scores 43.6% against 34.4% for its predecessor, on DeepSWE v1.1 it reaches 65.3% against 49.0%, and on AutomationBench it more than doubles, from 17.0% to 30.4%.

Google is pricing it to be used. Input costs $0.75 per million tokens and output $3.75 until the end of December, after which both double. It also powers Gemini Spark for AI Pro and Ultra subscribers across more than 160 countries.

Advertisement

The safety work is pitched at the frontier categories. Google says it added safeguards against cyber offence and chemical, biological, radiological and nuclear misuse, in line with its Frontier Safety commitments.

What stands out is the cadence. This lands three weeks after the last Flash release, and Google put out three Flash models in July alone, including a security-tuned variant.

The Pro line has not moved in that time. Google’s next Pro model is months behind schedule after its coding performance fell short of internal targets, which is awkward given coding is where the money is.

It may not arrive at all. Axios reports Google could skip 3.5 Pro entirely and go straight to Gemini 4 Pro, and the company declined to discuss the model’s fate.

Advertisement

Sundar Pichai told investors in July that Google intends to ship models faster and is already spending heavily on compute to train Gemini 4. Shipping the cheap tier every three weeks is one way to look faster while the expensive one stalls.

There is a cost inside the building. The delay has been hard on DeepMind morale, at a moment when rival labs are hiring.

Source link

Advertisement
Continue Reading

Trending

Copyright © 2025