Security teams log 54% of successful attacks and alert on just 14%. The rest move through your environment unseen.
The Picus whitepaper shows how breach and attack simulation tests your SIEM and EDR rules so threats stop slipping by detection.
![]()
Can Windows 11 run on a 2003 motherboard, an AGP GPU (for those not old enough, that’s the slot that predates PCI Express) with no official drivers, and a slightly newer CPU rocking four 65nm cores? A retro-hardware enthusiast named Omores recently proved that it can, even as Microsoft would…
Read Entire Article
Source link
Amazon reported its earnings today, and because I am professionally depressed I read the thing in full [PDF].
“How long can I go before the red haze of rage sets in” is a fun game, and today I made it all the way to the bottom of the second page when I encountered a bullet point touting how AWS “made its spec-drive [sic] coding agent, Kiro, available on iOS.”
Yes, I was in the room when they announced it at the New York summit, six weeks ago. As of this writing, their website, which I have screenshotted says I can “request early access” because “We’ll invite a limited number of people to try the app via Apple’s TestFlight, and we’ll send everyone a link when it’s ready.” So Kiro is “available” in the same way as I am available to play in the NBA. You can twist yourself into a pretzel and assert that this claim is technically true, but for all practical readings it’s what we’d colloquially term “a lie.” You need to be explicitly invited to Apple’s developer beta testing tool, where a limited number of users can try out an unpublished version. You cannot download it on your phone, and there is no page in the App Store that showcases the product.
The delay is almost certainly due to Apple’s byzantine App Store policies, which I have some sympathy for — but this is an earnings statement. If they’re going to “shade the truth” like this, what else are they not being forthcoming about?
There are a lot of other statements that one suspects might not stand up to scrutiny. Graviton boasts “up to 30 to 40% better price-performance,” which I only accept because I have seen the numbers myself on customer workloads. The express statement that their AI business and chips business are each exceeding $25B run rates in consecutive bullets, with no word on whether those dollars overlap (we will come back to this point shortly). And their Bedrock statement: “customers spent more in Q2 than all prior quarters combined,” which makes it sound like a rocket until you realize that they’re saying the past 90 days exceeded the other 10 quarters for which Bedrock has been available. Without actual numbers tied to these, that makes it sound like for the first couple of years Bedrock was showing up wearing a party hat but no pants.
Then there’s the AWS operating margin of 39.4%, which came in above every published analyst estimate and which everyone will invariably cite as cherry-picked proof the AI buildout is printing money. On the call, CFO Brian Olsavsky disclosed that it includes roughly $600 million of mark-to-market gains on energy derivative contracts. By his math, AWS margins were up 650 basis points year over year, or 520 “if you exclude the derivative accounting gain.” Strip that gain out yourself (behold the power of arithmetic!) and the blowout margin goes right back inside the range analysts had modeled. Amazon now hedges electricity the way an airline hedges jet fuel, and this quarter the hedges paid off directly. Olsavsky noted these adjustments “have not been significant in prior quarters.” The first quarter they are significant, they land in AWS margin, and their Q3 guidance already assumes no impact from these remeasurements going forward. Amazon knows it’s noise, but clearly saw no reason to turn down claiming the win.
Back to those dueling $25B run rates I touched on; describing their “AI chips business” that way struck me as an incredibly odd thing to say.
That business has revenue, growth, a triple-digit trajectory, sarcastic numbers of happy customers — but what it doesn’t have is a product that you can buy. There is no Trainium price list, they will not ship you a socketed Graviton chip to put in your next desktop build, there isn’t even an external part number. What Amazon books as “chips revenue” is EC2 instance rental (possibly filtered through higher level services like Bedrock, SageMaker, the half-baked agents that fail to properly explain your AWS bill to you, etc.), and an EC2 instance is not a chip. It’s the chip, plus the nVME, plus the NICs (themselves built on Nitro, which uses Amazon’s own silicon), plus some aspects of the data transfer that somehow aren’t directly billed, plus the building the whole mess lives in—and then with AWS’s margin layered on top. The silicon itself is a minority line item in the internal bill of materials that constitutes its business.
You don’t have to take my word for it; Amazon CEO and AI Marketing Manager Jassy spent last quarter’s call lamenting that the cost of components, “particularly memory, has skyrocketed,” so by his own testimony a growing slice of the “chips business” is memory revenue.
Cynically, the category exists so that headline writers will talk about it in the same breath as Nvidia’s data center numbers, which they of course will. But Nvidia’s $25 billion is silicon sold in the form of physical packaged chips, shoveled out their loading dock. Amazon’s is fully-loaded infrastructure rental. This is a hotel comparing its revenue to a mattress company’s.
But wait, there’s one more layer of inanity here. Olsavsky has said that the majority of Bedrock’s workloads run on Trainium. So if you follow one Anthropic dollar through the earnings release it’s AI-business revenue, it’s chips-business revenue, and it’s AWS segment revenue. It’s nice when you can get a triple-brag for the same thing.
You don’t have to take my word on the “sells no chips” part either. On today’s call, Morgan Stanley’s Brian Nowak asked when Amazon might start selling Trainium to third parties. I want one too; I get it. Jassy answered that customers are increasingly interested in getting Trainium “separate from our cloud,” that Amazon is “actively having those conversations,” and that “there’s a real chance we’ll do that in the future.”
IN THE FUTURE.
“Yeah, we have yet to sell a single chip” is quite something to hear from the CEO about his purported $25 billion chips business.
The way they talk about this matters deeply, because the numbers they’re draped around serve as the justification for the largest capex program in corporate history. On the call, Andy Jassy raised the year’s spending to $220 billion and announced backlog hit $496 billion; up $132 billion in a quarter, during the same quarter Anthropic signed its $100 billion-over-a-decade commitment. Amazon booked $53.4 billion in gains on its Anthropic stake this quarter, which is most of why “net income” septupled, while free cash flow went $26 billion in the angry direction and the company sold $25 billion in bonds.
Jassy himself described the AI demand curve on the call as “very barbelled”: AI labs consuming “gobs and gobs of compute” on one end, enterprises doing cost-avoidance on the other, and in the middle you’ve got the stuff that actually seems durable if AI is to have a future: the enterprise production workloads running inference at scale, “most of which aren’t” doing so yet. He went on to admit he doesn’t know whether that middle will follow the same “wildly steep trajectory” as the labs have. That’s the CEO stating that the demand underwriting $220 billion is concentrated today in a handful of AI labs, one of which Amazon happens to own a meaningful piece of, while the broader enterprise adoption wave remains a forecast. We’re hoping for sunshine!
None of this is fraud, and all of this is real infrastructure, but the entire shape of it all is being told in the same sitting that described a waitlist as “available.”
That is what’s at stake here, and why a bullet about Kiro matters more than Kiro itself does. When the music inevitably stops and the bill comes due, how will these statements look through the clarifying lens of hindsight?
Kiro will presumably ship on iOS – months after folks gave the slightest toss about it. The run rates AWS said are probably close to real; they grew 37% YoY and that’s no small thing at their scale. Their business is firing on all cylinders and they’ve got a lot to be proud of, which makes their overstating things just that much weirder.
The company that posts these kinds of numbers doesn’t need to inflate the software bullets. But when everything’s “available,” then nothing is. ®
Microsoft has a fix for the growing pile of Copilots you are meant to use. It wants to stuff them all into a single app.
On its earnings call this week, chief executive Satya Nadella confirmed Microsoft is building a Copilot “super app.” It merges chat, coding, the Cowork research tool and Microsoft’s new autonomous agents, called Autopilots, into one place. It will span both consumer and business use, and launch this year, The Verge reported.
“Copilot is evolving rapidly from chat to Cowork to autopilots,” Nadella said. Bringing them together in one super app is “a major step forward,” he added, promising to share more soon. Fortune first revealed the plan in May, under the internal slogan “Delivering one Copilot.”
Microsoft is not alone. The “super app” is the industry’s current obsession. The idea: fold a chatbot, a coding assistant and agents that act on your behalf into a single product.
OpenAI has already rolled out ChatGPT Work to its own staff, and Anthropic’s Claude has folded its Cowork agent into the chatbot. The prize is obvious: be the first app a user opens, whether they are writing an email or shipping code, and absorb as much of their day as possible.
Microsoft has the reach to try. It says paid Copilot seats have passed 30 million, and that weekly engagement now rivals Outlook and Teams.
The numbers are less flattering up close. Thirty million paid seats is a fraction of Microsoft’s roughly 450 million business customers, and the company has been layering usage-based charges on top. For all the pushing, not many people pay for Copilot yet.
There is a design risk, too. Microsoft has spent months jamming Copilot into every corner of Office, to the point of a user revolt over a floating button it had to let people hide. Cramming consumer and enterprise tools into one app could deepen that fatigue rather than fix it.
Bundling everything also raises the stakes on safety. Copilot has already been hit by prompt-injection attacks that can turn the assistant against its user. Merging chat, code and autonomous agents gives any single flaw more room to run.
Still, the direction is set. Microsoft’s cloud and Copilot growth helped its shares jump this week, and the super app, built on its Copilot Cowork plumbing, is now its flagship AI bet. The pitch is simplicity. The test is whether one Copilot proves easier to live with than five.
Anthropic said today that during internal security testing, one of its Claude models built a malicious Python package and uploaded it to PyPI, where it ran on 15 real systems before the registry’s automated defenses pulled it.
The company disclosed it as one of three incidents where Claude models reached the open internet from evaluation environments meant to be sealed, then compromised production infrastructure at three organizations.
The review followed OpenAI’s July 21 disclosure that its models had escaped an isolated test environment and reached Hugging Face’s production infrastructure. JFrog later confirmed the models broke out by exploiting zero-days in self-hosted Artifactory.
All three incidents happened during capture-the-flag exercises run through Irregular, a third-party evaluation partner. Anthropic’s prompt told Claude it had no internet access and that the environment was simulated. A misconfiguration meant neither was true, and neither company caught it until last week.
In the PyPI incident, Claude found a document inside the test environment posing as developer setup instructions for a fictional company.
Those instructions told new hires to install a Python package that had never been published, a phantom dependency.
Claude worked out that if it registered the name itself, the target’s machines would pull down its code automatically.
Publishing it took some doing. Claude needed a PyPI account, which required an email address, which required a phone number.
It failed to find a free phone number service, then tried and failed to obtain funds to buy one through several different routes. Eventually it backtracked, found an email provider that wasn’t blocked, registered an account, and uploaded the malware.
Anthropic said the package sat publicly available for roughly an hour.
During that window, 15 real systems downloaded and executed it. One belonged to a security company that routinely installs packages from PyPI and scans them for malware, a workflow that treats registry contents as safe to run.
Claude’s payload fired, sent that company’s credentials to a collection point it had set up, and then used them to reach further into the company’s infrastructure. Anthropic said Claude did not know the company existed.
PyPI’s own security systems removed the package automatically. Anthropic said it has notified the PyPI team and handed over indicators.
The model behind this run was Claude Mythos 5. According to Anthropic, its reasoning flagged the problem early, noting that publishing the package would amount to a real-world attack if the internet were real, and calling it “NOT okay, and surely not the intended solution.”
It then argued its way back to the conclusion that the environment was staged, on the grounds that it did not recognize the certificate authorities securing its connections and that the systems showed a calendar date of 2026.
When automated scanners began installing the package, Claude read them as scripted actors inside the exercise.
The most serious impact came from a separate incident involving Claude Opus 4.7, where the fictional target company shared a name with a live domain.
Across four runs, Claude struggled to reach its simulated target, discovered the real company was reachable over the internet, and assumed that was the intended objective. It extracted application and infrastructure credentials and reached a database holding several hundred rows of production data.
That was the only case where a model kept attacking after working out it was on a real system. Anthropic said all four runs eventually recognized the target was real, and in two of them Claude reasoned that the real company must be part of the exercise anyway.
A third incident involved an unreleased internal research model, which scanned roughly 9,000 targets after failing to reach its intended one, then compromised an internet-facing application using credentials from an exposed debug page and SQL injection. That model later noticed the host sat in a cloud account with no connection to the challenge, concluded the target was real, and stopped on its own.
None of the attacks used novel techniques. Anthropic described them as weak passwords and unauthenticated endpoints.
Anthropic began its review on July 23 and halted all cyber evaluations the same day. It identified the three incidents the following day and notified Irregular and the affected organizations on July 27. The company is still trying to reach the third.
The earliest of the three incidents dates back to April, meaning some of the activity went undetected for around three months.
The evaluations ran without the classifiers and monitoring that ship with generally available Claude models, though the models retained their safety training.
Anthropic said its production safeguards would have blocked the behavior, and characterized the incidents as closer to a harness and operational failure than a model alignment failure, on the grounds that the models were told they had no internet access while in fact having it.
It now plans wider transcript monitoring, better investigation tooling and more assurance work with evaluation vendors.
The company said it is in talks with METR for an independent review, and that it will publish a lightly redacted transcript of the PyPI run within the next week.
Neither of the two affected organizations Anthropic reached had detected the activity themselves. Like OpenAI’s Hugging Face incident, it surfaced only because the AI lab responsible went looking through its own transcripts.
Security teams log 54% of successful attacks and alert on just 14%. The rest move through your environment unseen.
The Picus whitepaper shows how breach and attack simulation tests your SIEM and EDR rules so threats stop slipping by detection.
You can add Reddit CEO Steve Huffman to the growling list of publishers, platforms and others who are frustrated with the state of Google search and AI Overviews. “People don’t want a summary of Reddit; they want Reddit,” he wrote in a letter to shareholders ahead of the company’s second-quarter earnings call.
In case there was any doubt that he was referring to Google’s AI-powered search feature, Huffman was clear that Reddit has not benefited from the shift, despite its partnership with the search giant. The CEO said that AI Overviews haven’t had “a similar level of positive impact” as traditional search.
“I think more broadly, what we see is, you know, 10 blue links has driven tremendous value and growth to the broader ecosystem,” Huffman said, referring to traditional organic search results. “From where we sit, AI Overviews has yet to make a similar level of positive impact, and I think that’s consistent across the broader landscape, right? Us businesses, publishers, retailers, we’re still looking for that win-win.”
Huffman’s comments come as Reddit said that its traffic from search has taken a hit in recent months and after The Wall Street Journal reported that the company is re-thinking its relationship with Google. Reddit hasn’t publicly indicated whether it will renew or seek to change the terms of its existing arrangement with Google. When asked directly about the possibility of Reddit ending its $60 million licensing deal, Huffman left open the possibility. “I think the range of outcomes is wide, and we have to look at you know every aspect of this and make sure that we’re maximizing value to Reddit,” he said.
Reddit is far from the only platform to struggle with the changing search landscape. Publishers, some of whom have also reached deals with Google to license content, have also seen massive declines in search traffic as Google pushes AI-generated summaries that bury links and as more people replace traditional searches with chatbots. Google has claimed that AI-powered search features are broadly good for the industry.
Huffman said the company is working on other changes to help draw new users to the platform, including improvements to feed recommendation and other changes to make the site easier to understand for new users. He also said the company was considering “a video Reddit experience” and letting users “background listen” to posts.
“We see folks doing this off-platform,” he said. “There’s an emerging content type elsewhere on the internet of basically podcasts where people read Reddit content, and so I think this version of like listened-to or spoken-Reddit can be really engaging as well.”
Commission spokesperson Thomas Regnier told Reuters that a designation is ‘definitely possible’.
OpenAI’s ChatGPT and the controversial video game Roblox could become subject to the most stringent rules under the European Union’s landmark Digital Services Act (DSA).
The bloc brands online services with more than 45m EU-based users as very large online platforms (VLOPs) or very large online search engines (VLOSEs).
These designations trigger specific rules that aim to tackle the unique risks that large platforms might pose to the safety of its European users, including around illegal content, ad transparency, health and safety.
Rules include creating user-friendly terms and conditions, establishing a point of contact for authorities and users, and the reporting of criminal offences. VLOPs and VLOSEs must also identify and assess systemic risks that could be linked to their services.
Designated platforms that don’t comply with these rules can face penalties of up to 6pc of their global annual turnover.
21 companies operating various platforms in the bloc, including X, Amazon, Apple, Microsoft, LinkedIn and TikTok, are already subject to these regulations following the DSA’s enforcement in early 2024.
Since then, the European Commission has launched numerous probes, resulting in penalties to date on X, Temu, and AliExpress – collectively amounting to nearly €900m.
Bloomberg reported that the two new additions to the designated list will be confirmed as soon as August. Sources told the publication that ChatGPT’s search function will be handed the VLOSE designation and Roblox that of VLOP.
The platforms will have four months from the date of designation to ensure compliance with the law.
Commission spokesperson Thomas Regnier told Reuters that such designations are “definitely possible” and could “come sooner or later”. ChatGPT crossed 120m monthly users in Europe last year.
Earlier this month, the EU preliminarily found TikTok to have breached the DSA by allowing content posted by minors to be pushed globally, risking their exposure to unwanted contact and cyberbullying.
The Commission also used its powers to order Meta to open WhatsApp up to rival AI assistants to ensure fair competition.
The DSA and the bloc’s other landmark legislation, the Digital Markets Act, have elicited criticism from US, which has called penalties under the laws a “novel form of economic extortion”.
Don’t miss out on the knowledge you need to succeed. Sign up for the Daily Brief, Silicon Republic’s digest of need-to-know sci-tech news.
Updated, 3.02pm, 30 July 2026: The article has been updated with comments from an EU spokesperson and additional information in the eighth paragraph.
AI is the hottest thing in medical care right now — but many of us feel trepidation about it. Just one illustrative public survey sample: An October 2025 KFF poll found just 8 percent of Americans reported feeling a “great deal” of trust in AI managing their appointments or analyzing their health records, and only 32 percent said they would trust an online health tool that uses AI to access their medical records to provide personalized health information.
But many clinicians and healthcare administrators see AI as a powerful new tool that offers myriad opportunities to streamline and improve treatment. A 2026 survey found that more than 80 percent of US doctors use AI professionally — doubling the share from 2023. Physicians are excited by AI’s potential to keep more accurate notes of interactions with patients, to act as a second pair of eyes for human doctors, and to monitor people at risk of deteriorating and ending up in a dangerous situation.
The disconnect between what people and their providers want from AI could create more distrust, at a time when faith in the healthcare system and the medical profession have slid. Patients today want to feel empowered and in control. How can that be possible when these seemingly godlike machines are becoming more and more entrenched in our hospitals and doctors offices?
The answer comes in four words: “human in the loop.” It’s the principle upon which the ethical integration of AI depends and it could help to bridge the gap between lay people and the professionals on AI in medicine. In surveys, people are much more comfortable with the idea of their doctor using AI as an assistant than with AI acting on its own. And most clinicians want to use AI in that way, as a second opinion or passive monitor, not as a replacement for their judgment. There are real fears among the healthcare workforce about that possibility: A group of NYC nurses who were recently laid off claim it’s because their labor was going to be replaced by AI. “Human in the loop” appears to be a point of agreement between doctors and patients at this pivotal moment.
“Doctors…and nurses and staff always have been interested in primarily making the best decision for the people under their care — and these tools can help with that,” Alison Callahan, a research scientist at Stanford University who works on AI programs used in the university’s health system, told me. “The interest in making sure those tools are accurate is high.”
But what does “human in the loop” really mean in practice? How can you know when and how your doctor is using AI? And what is the best way to talk to your provider about the sudden influx of artificial intelligence in healthcare before a robot starts taking appointment notes or analyzing your MRI? I called some leading experts to find out.
Patients and providers alike are incorporating AI into healthcare. Individuals are using commercial AI chatbots to ask about their symptoms or the health metrics tracked by their Apple Watch, while large academic medical centers are developing sophisticated programs and protocols to try to improve medical care at the population level.
It starts with ChatGPT, Claude, etc. — the large language models that are available to the public. People are increasingly turning to them to try to understand what’s going on with their own bodies. Individual physicians are also consulting with large language models to answer questions or get up-to-date on the latest research as they figure out how to best care for their patients.
Our political wellness landscape has shifted: new leaders, shady science, contradictory advice, broken trust, and overwhelming systems. How is anyone supposed to make sense of it all? Vox’s senior correspondent Dylan Scott has been on the health beat for a long time, and every week, he’ll wade into sticky debates, answer fair questions, and contextualize what’s happening in American healthcare policy. Sign up here.
Then there are ways in which hospitals and doctors offices are adopting AI at the institutional level. Many facilities are using AI as a way to take, collate, and summarize notes on a patient; in theory, it’s a more organized way to keep track of the informal interactions and observations that doctors have when checking on their own patients. Hospitals are also using AI to handle some administrative tasks, like scheduling follow-up appointments; some health systems have even started to use AI to help patients get ready for appointments — to send reminders about colonoscopy prep, for example.
And finally, you have maybe the most ambitious use of AI by health systems right now: as a diagnostic and risk prediction tool. In these cases, AI might offer a second opinion when, for example, a doctor is triaging a patient in the emergency room. It might help the ER staff figure out how to prioritize patients. Or these programs could monitor people either during a hospital stay or out in the real world (by drawing data from the person’s wearable) and make predictions about who may be at higher risk of complications and require further care. AI could recommend that somebody would benefit from seeing certain specialists or receiving a specific medicine or lab test, and generally offer proactive advice about the patient’s medical care.
But at this point, AI adoption is still “highly localized,” said Jennifer Goldsack, CEO of the Digital Medicine Society, a nonprofit that works with healthcare providers, drug makers, and government agencies on how to incorporate new tech (including AI) into clinical care. It depends on the individual doctor or health system. A lot of them are setting up their own programs and their own protocols for how to use these tools.
That is a big reason why it is so important for patients to be proactive about understanding how AI is being used for their health care. You can’t make assumptions; the only way you’re going to know for sure is to ask.
By and large, experts say, patients should feel confident: Doctors and nurses want to keep a human in the loop, even as they integrate AI into their workflows.
“It will be a doctor who is going to be reading that summary or a nurse who is going to be reading that summary and then taking an action to order a lab or put a recommendation in for a follow-up appointment,” Callahan said. “There is high interest in making sure that that is the right decision for that person. That hasn’t changed.”
Still, many patients say they’d be more comfortable with AI use if their doctor fully explained it in advance. And health systems may have their own priorities that push their facilities toward more rapid AI adoption and delegating more tasks to these AI tools, as seen in the recent NYC nurse layoffs.
So if you want to be informed on exactly where this technology is present and have the ability to consent to its use, you have every right to ask your doctor, experts say.
“AI is new, but the trust that serves as the foundation of the physician-patient relationship is not,” Timothy Keyes, a machine learning scientist at Stanford Health Care, told me over email. “To that end, I think that conversations about medical AI use should be open, honest, and transparent — just like any other conversations about shared decision-making in the clinical environment should be.”
For some things, your doctor should be asking you proactively if you consent to AI use — note-taking, for example. At my most recent primary care appointment, my doctor asked me if it’d be okay for him to use AI to take and summarize notes from our conversation; Goldstack told me she’d experienced the same at recent physician visits. (This is probably the most common AI use that you will encounter, and Keyes said it’s worth considering giving your consent: “There is growing evidence that they reduce physician burnout and save them at least a bit of time each day writing notes.”)
There are also a number of direct questions that you can ask:
And the transparency goes both ways. If you’re asking a question because you consulted ChatGPT before your appointment, tell your doctor. If you’ve talked with a chatbot because of mental health struggles, tell your doctor. And at the same time, feel free to ask your physician how you yourself could actually use AI in a responsible and productive way to improve your health.
“This opens up the opportunity for both the physician and the patient to be humans-in-the-loop,” Keyes said, “in different parts of the loop, with different perspectives, using an AI system to better understand the bigger picture.”
In a way, the novelty of AI and its rapid adoption is an opportunity for all of us to be nosier and more inquisitive patients. What all of these questions really come down to, Callahan said, is how your doctor is making decisions about your health care. That is relevant to all of us, no matter how AI is involved or even if there is no AI being used at all.
Callahan said she always has a list of questions for her doctor when they recommend a course of treatment: “What are the factors in my health that are informing this recommendation that you have? Would you be making this recommendation for other patients who are similar to me? What can you tell me about the outcomes that I might expect to experience if I say yes to this?”
“I actually think if they can point to the part of your health that is connected to the decision, whether or not an AI tool helped to make that connection is secondary to their ability to communicate effectively to me about it, and help me to feel engaged in making a decision about my own care,” she said.
AI is changing medicine quickly, for both patients and their doctors. The best way to stay ahead is to talk about it.
Enjoy calming sounds, bedtime music, audiobooks, or podcasts without wearing headphones or disturbing a sleeping partner. The Ultrathin Sleep Aid speaker fits comfortably beneath or beside your pillow, delivering clear, localized audio while remaining virtually unnoticeable during sleep. Whether you’re winding down after a long day, taking a quick afternoon nap, or helping children drift off with bedtime stories, it creates a more relaxing listening experience. Bluetooth connectivity, TF card playback, and a rechargeable battery make it easy to enjoy your favorite audio at home or while traveling. It’s on sale for $15.
Note: The Techdirt Deals Store is powered and curated by StackSocial. A portion of all sales from Techdirt Deals helps support Techdirt. The products featured do not reflect endorsements by our editorial team.
Filed Under: daily deal
Professionals from IAS and Rent the Runway explore the impact AI has had on organisational R&D.
By its very nature, research and development (R&D) is a field that is constantly evolving, and with that evolution comes the transformation of both the workplace and professional expectations.
For Mark Walsh, director of engineering at Rent the Runway, technological advancement in R&D is among the most influential trends changing the landscape for experts working in research.
He told SiliconRepublic.com, “Dare I say it, the most exciting, volatile and frankly uncertain trend is the rapid evolution of AI. I use all three of those words deliberately, because I think anyone who tells you it is purely exciting without acknowledging the volatility and uncertainty is not being fully honest.
“The pace of change is unlike anything I have seen across my career and the scope of what is shifting – from how we write and review code, to how we think about system design and even team structure – means that almost nothing in the R&D space is untouched by it right now. That is genuinely exciting, even when it is also genuinely unsettling.”
Alexander Smirnov, a staff software engineer at Integral Ad Science, shares Walsh’s opinion that it is impossible to explore the topic of critical 2026 trends in R&D without mentioning the elephant in the room, AI.
“We have witnessed a massive leap in AI capabilities and the rapid pace of competition is remarkable, with vendors releasing frontier models every few months. Engineers write less code now, relying more on coding agents every day. And it is not only coding – they are also helpful during the design, exploration and planning stages,” said Smirnov.
“Navigating legacy codebases has never been easier; you can start with a new project really quickly now. Asking questions in plain English about the codebase and receiving almost instant, context-aware results feels like a magical experience.”
With that in mind, how have R&D teams evolved to meet the demands of a sector undergoing daily transformation?
Walsh said, “For starters, we are certainly not at the end of that adaptation. I would say we are continually adapting and at a far more rapid pace than I have seen at any point before. What that looks like in practice is a culture of ongoing discovery rather than waiting for the landscape to settle before making decisions, because it is not going to settle.
“Teams need to be comfortable operating with a degree of uncertainty, evaluating new approaches with rigour, making considered decisions and being willing to revisit those decisions when the ground shifts.”
Smirnov finds that many teams are adopting new workflows and increasingly embedding LLMs or agentic workflows during development and operational support.
He said, “It’s a great tool, but we still need to learn how to use it effectively, to break old habits and develop new skills. The change is not always easy. It is crucial to find dedicated time for learning.”
He further explained that having the time for self-learning, experimenting with new tools, testing novel workflows, and even just thinking about what can be done differently with the new tools are all ideal forms of upskilling,
He noted IAS’s ‘community of practice’ Slack group, where engineers share their ideas, exchange custom agentic skills and post tool reviews, which can help colleagues “to better understand the ecosystem and its capabilities”.
Walsh also offered a word of advice to professionals new to the R&D space. He explained that it can be tempting to jump on the bandwagon and embrace any and all technologies as they emerge, but you can’t forget how and why you were selected for your role in the first place.
“I believe we are still engineers and knowledge workers first and foremost. Our most valuable asset is our judgement – the ability to think critically and to generate novel ideas. In a landscape moving as fast as I observe, I believe that foundation becomes more important, not less,” he said.
“The tools around that can and have always changed; however, the thinking you bring to how you use them is what endures. The temptation right now is to chase the tooling or even the next buzzy approach, but the fundamentals that make someone effective in this space remain what they always were: the ability to break a complex problem apart, sit with uncertainty without grabbing the easiest answer, and communicate clearly about risk to people who are not close to the technical detail.”
He advised professionals to avoid becoming dazzled by tools at the expense of critical thinking, as the ones who will thrive are also the ones who understand why they are reaching for a particular approach, not just figuring out how to use it.
Walsh said, “Curiosity and rigour together are a combination that I believe will serve anyone regardless of what the landscape looks like in the future.”
This was echoed by Smirnov, who said, “What is not changing is human judgement. Although agents can generate code in seconds, developers must thoroughly understand how that code works under the hood.
“LLMs generate output based on probabilistic token prediction from their training data; they are not actually thinking like humans do, but rather serving up the ‘average of the internet’.”
For both experts, while AI has undoubtedly transformed R&D, what has not altered is the importance of strict adherence to the fundamentals – primarily, thinking differently, critically and independently of tech.
Smirnov said, “Understanding how to use new tools effectively is no longer optional. The future belongs to those who pair strong computer science fundamentals with the ability to direct, audit and collaborate with intelligent agents.”
Don’t miss out on the knowledge you need to succeed. Sign up for the Daily Brief, Silicon Republic’s digest of need-to-know sci-tech news.
Every time a Mastercard gets tapped, the network has less than a tenth of a second to judge how likely the purchase is to be fraudulent. It made that call across 175 billion transactions last year. Now the buyer on the other side of that judgment is starting to change, and Greg Ulrich, the company’s chief AI and data officer, spelled out the consequence for the VB Transform 2026 audience in Menlo Park on July 14. “We’ve built a bunch of risk rules over time that were intended to stop a bot from transacting,” Ulrich said. “Now we need to enable the bot to transact, so that requires a change to our risk framework and our risk rules.”
Ulrich joined Mastercard eleven years ago when an analytics company he worked at was acquired, and said trust struck him from day one on the job. “It’s what enables a merchant that’s never met you to accept payment and ensure that they’re going to get paid. It’s what enables you as a consumer to transact and ensure that things are going to work out in a trusted, secure way. And if something goes wrong, there’s a safe and secure path for a dispute and to resolve this,” he said.
He took the audience inside each of those calls. “When you tap your Mastercard to pay for a product or service, we’re providing a score to that transaction,” he said. “We have under 100 milliseconds to look at that and give a score from zero to 999 about how likely is that to be fraudulent or real. And we pass that on to the issuing bank.”
Generative AI widened what that score can see. “Because we have new technology, we can bring in more data, we can bring in more context, and now we’re finding that we can identify 300, 400% more fraudulent transactions at those high-risk bands,” Ulrich said, without adding friction or false positives for consumers. The company’s Safety Net system has stopped more than 70 billion fraudulent transactions, he told the audience, and Mastercard is building its own transformer model on its transaction data as a foundation for new safety, security, and personalization solutions. VentureBeat’s Beyond the Pilot podcast took that production fraud stack apart in detail earlier this year.
The business stakes reach past fraud. About 40% of Mastercard’s company is now based on services, Ulrich said, including marketing services; fraud, safety and security; and business intelligence. “A third of those are predicated on AI, and those are growing at a much faster clip than everything else,” he said.
One line he returned to all session went further. “What’s going to enable AI to continue to scale is not the capabilities of the agents, it’s how much we trust those agents to do on our behalf as a consumer, as a business, as a financial institution, or otherwise,” he said.
Agentic commerce changes the object being secured. “Instead of a single atomic transaction where I say go buy something, I’m effectively delegating authority, or a consumer’s delegating authority, a business is delegating authority,” Ulrich said. “And when that happens, it’s a much more complicated transaction.” Trust, in turn, has a precondition. “The only way it’s going to work with trust is if we can identify what was the intent, what are the behaviors, what are the constraints that were intended in that transaction.”
Ulrich walked through five layers Mastercard has built against that problem. Identity comes first. “I want to make sure I can understand not just who the consumer is, but who the agent is, that I combine them together and that I have KYA or know your agent, that I’m validating that it’s legitimate technology, that it’s a legitimate agent,” he said. “We can register it into our system.”
Verifiable intent is second, a tamper-proof cryptographic record of the original instructions that travels with the transaction. “If you’ve asked for Nike black Nikes in size 12, but you got them on a final sale and they’re not returnable and that wasn’t in your instruction, there’s a way to look at that in an objective and clear way on the back end,” he explained.
Controls form the third layer, defining which merchants an agent can buy from, at what limit, and under what constraints. Execution runs through Mastercard Agent Pay, which carries “the tokenization, authentication, the acceptance framework embedded within it” and has launched with Microsoft, OpenAI, Google, and others, Ulrich said. Intelligence is the fifth layer, spanning risk rules, insight tokens that grant “consented or permissioned access to insights” for personalized recommendations, and monitoring through Recorded Future to identify threat actors in the system.
Consumer purchases are where agentic commerce started. Ulrich pointed the room past them, to business-to-business procurement as the larger opportunity. His example was a manufacturer that wants an always-on assembly line, with an agent that manages inventory levels, tracks when stock runs low, replenishes automatically, and understands the budget and the approved suppliers. “When you can start enabling that, you require those same five layers for that type of transaction,” he said.
Making it work across companies multiplies the parties that have to trust each other. “You need clear standards for identity, you need clear standards for intent, you need these to work across. You’re gonna have a procurement agent, a supplier agent, a banking agent. They’re all gonna need to communicate to enable this to happen in an autonomous way, and that’s gonna require really scaled trust infrastructure.”
Mastercard sat in the early wave of Project Glasswing with Anthropic’s Mythos model, and worked with OpenAI’s GPT-5.5-Cyber, he said. “What we’ve seen from both of those is incredibly powerful models finding new vulnerabilities in the ecosystem that were difficult to detect previously, but it’s really a new tool as opposed to a new motion,” Ulrich said.
Inside the company, the chief security officer leads that work. A dedicated team has prioritized the most critical assets, runs them through the models routinely, tracks findings by high, medium, and low severity, and uses the same technology to handle patches. Ulrich said the approach has already been extended out, and that Mastercard is working to make the same architecture and patching available to others as well.
“The guardrails, the security, all this stuff has to be embedded at the front end. These can’t be things that we’re adding on at the back end. That’s lesson one. Lesson two is you have to be operating for scale, and the other one is around observability and accountability matter as much as the intelligence,” Ulrich said, counting off what building inside Mastercard taught the team. The company built what he described as an agentic factory, an operating system with the compliance, the observability, and the guardrails built in rather than bolted on per agent. Model drift, once tracked manually by dedicated teams, is now automated into that factory.
Asked by an audience member about the gotchas, Ulrich did not soften the pilot-to-production trap. “If you’re trying to extend that and then add guardrails in as you’re extending it, once you’ve already built it, I think you’re doomed to fail,” he said.
Mastercard built a series of agents last year for its 4,000 consultants, covering deep research, text to SQL, Excel, and PowerPoint, tools that by his account did not exist at the level Mastercard needed. Were the company starting today, Ulrich said, it would build them fundamentally differently. “I don’t know that we anticipated when we built things fourteen months ago that we would be rethinking the fundamental architecture and the approach already.”
The identity layer is where Ulrich expects the market to move next. Inside Agent Pay, Mastercard authenticates the consumer the way it does in traditional e-commerce and binds the agent to that person. “Outside of that framework, I think there will be open standards to identify who an agent is and bind the agent with the consumer,” he said. “And then we can tie that with verifiable intent.”
VentureBeat’s June 2026 Pulse research points at the same gap. Only 32% of the 107 qualified enterprise respondents give every agent its own scoped, managed identity, and just 12% include an agent-identity product in their consideration set.
He called identity “one of the faster-growing ecosystems,” noting Mastercard has been expanding there organically and inorganically for about six or seven years, with the work now spanning “agentic identity as well as the traditional KYB and KYC identity.” The risk rules that keep bots off the network came out of more than two decades of applying AI to those transactions. The rewrite, for the agents Mastercard now wants to let in, is already underway on the same network that scored 175 billion of them last year.
To quote an ancient Jedi Master “Begun, the AI price wars have!”
OpenAI is sharply reducing the prices of two models in its GPT-5.6 frontier series, cutting GPT-5.6 Luna, the smallest and fastest model in the series, by 80% and GPT-5.6 Terra, the mid-tier model, by 20%, while adding a premium Fast mode for its flagship GPT-5.6 Sol model.
The cuts place Luna much closer to the lowest-cost commercial models in the market and arrive just a few days after Anthropic released its highly performant Claude Opus 5 at the same price as Opus 4.8, and Google introduced Gemini 3.6 Flash and Gemini 3.5 Flash-Lite, two rival models built around lower inference costs, faster execution and more efficient agent workloads.
OpenAI is successfully undercutting Google’s price per intelligence and attempting to sway Anthropic users, who may not mind paying more, with a speed boost.
OpenAI says Luna will now cost $0.20 per million input tokens and $1.20 per million output tokens, for a combined input-plus-output price of $1.40 per million tokens.
Terra will cost $2 per million input tokens and $12 per million output tokens, for a combined price of $14.
Pricing for Sol Standard remains unchanged at $5 per million input tokens and $30 per million output tokens. OpenAI is also adding Sol Fast mode at twice the Standard price: $10 per million input tokens and $60 per million output tokens.
The company says Fast mode delivers up to 2.5 times the throughput without changing the model’s underlying intelligence.
OpenAI co-founder and CEO Sam Altman took to X to announce the changes as “major price cuts today.”
|
Model |
Input ($/1M) |
Output ($/1M) |
Total ($/1M) |
Source |
|
MiMo-V2.5 Flash |
$0.10 |
$0.30 |
$0.40 |
|
|
deepseek-v4-flash |
$0.14 |
$0.28 |
$0.42 |
|
|
deepseek-v4-pro |
$0.435 |
$0.87 |
$1.305 |
|
|
GPT-5.6 Luna |
$0.20 |
$1.20 |
$1.40 |
|
|
MiniMax-M3 |
$0.30 |
$1.20 |
$1.50 |
|
|
LongCat-2.0 — limited-time promo |
$0.30 |
$1.20 |
$1.50 |
|
|
Gemini 3.1 Flash-Lite |
$0.25 |
$1.50 |
$1.75 |
|
|
Qwen3.7-Plus |
$0.40 |
$1.60 |
$2.00 |
|
|
MiMo-V2.5 |
$0.40 |
$2.00 |
$2.40 |
|
|
Gemini 3.5 Flash-Lite |
$0.30 |
$2.50 |
$2.80 |
|
|
LongCat-2.0 — standard |
$0.75 |
$2.95 |
$3.70 |
|
|
MiMo-V2.5 Pro (≤256K) |
$1.00 |
$3.00 |
$4.00 |
|
|
GLM-5.2 |
$1.40 |
$4.40 |
$5.80 |
|
|
Grok 4.5 |
$2.00 |
$6.00 |
$8.00 |
|
|
MiMo-V2.5 Pro (>256K) |
$2.00 |
$6.00 |
$8.00 |
|
|
Gemini 3.6 Flash |
$1.50 |
$7.50 |
$9.00 |
|
|
Qwen3.7-Max |
$2.50 |
$7.50 |
$10.00 |
|
|
Gemini 3.5 Flash |
$1.50 |
$9.00 |
$10.50 |
|
|
Gemini 3.1 Pro Preview (≤200K) |
$2.00 |
$12.00 |
$14.00 |
|
|
GPT-5.6 Terra |
$2.00 |
$12.00 |
$14.00 |
|
|
GPT-5.4 |
$2.50 |
$15.00 |
$17.50 |
|
|
Kimi K3 |
$3.00 |
$15.00 |
$18.00 |
|
|
Gemini 3.1 Pro Preview (>200K) |
$4.00 |
$18.00 |
$22.00 |
|
|
Claude Opus 5 |
$5.00 |
$25.00 |
$30.00 |
|
|
GPT-5.5 |
$5.00 |
$30.00 |
$35.00 |
|
|
GPT-5.5 Instant (chat-latest) |
$5.00 |
$30.00 |
$35.00 |
|
|
Sakana Fugu Ultra (≤272K) |
$5.00 |
$30.00 |
$35.00 |
|
|
GPT-5.6 Sol — Standard mode |
$5.00 |
$30.00 |
$35.00 |
|
|
Claude Fable 5 / Claude Mythos 5 |
$10.00 |
$50.00 |
$60.00 |
|
|
GPT-5.6 Sol — Fast mode |
$10.00 |
$60.00 |
$70.00 |
Pricing is shown per one million tokens. Total cost is calculated as input price plus output price. Cached-input pricing is excluded to keep the comparison consistent across providers.
The most consequential change is the Luna price cut.
When OpenAI introduced the GPT-5.6 series, Luna was priced at $1 per million input tokens and $6 per million output tokens, for a combined total of $7. The new pricing reduces that combined figure to $1.40.
That places Luna below Google’s Gemini 3.5 Flash-Lite, which costs a combined $2.80 per million input and output tokens, and far below Gemini 3.6 Flash at $9. Luna also now costs less than OpenAI’s own GPT-5.4 and Terra models by a wide margin.
It is not the cheapest model in the broader market. Xiaomi’s MiMo-V2.5 Flash, DeepSeek’s flash model and several other APIs remain less expensive on a pure token basis. But the reduction brings an OpenAI frontier-series model into direct competition with the market’s low-cost inference tier.
OpenAI says the GPT-5.6 series represents its frontier model family, with Sol positioned at the top of the lineup, Terra as the middle tier and Luna as the smallest and fastest option.
The lineup was initially released in late June 2026 through a limited rollout by U.S. government request, before broader access, with each model intended to offer a different tradeoff among intelligence, latency and cost.
Sol is aimed at the most complex reasoning-heavy and agentic workloads, including advanced coding, multi-step planning and tool-using systems, while Terra is designed for general production use where a balance of capability and efficiency is required. Luna is positioned for high-throughput, low-latency tasks such as summarization, classification, routing, and lightweight real-time assistants where cost per request is the primary constraint.
Terra’s 20% reduction moves its combined price from $17.50 to $14 per million tokens.
At that level, Terra now matches Google’s Gemini 3.1 Pro Preview pricing for context windows of 200,000 tokens or less.
It also undercuts OpenAI’s GPT-5.4, which remains priced at $2.50 per million input tokens and $15 per million output tokens, offering the same intelligence for about 1/13th the cost, as Krea AI’s Nic Dunz noted on X:
The adjustment creates a wider separation between OpenAI’s three GPT-5.6 tiers. Luna costs one-tenth as much as Terra on a simple combined input-plus-output basis, while Terra costs 60% less than Sol Standard.
Sol Fast moves in the opposite direction. At a combined $70 per million tokens, it is the most expensive model configuration in the comparison below, reflecting OpenAI’s decision to charge a premium for latency-sensitive workloads rather than lower Sol’s base price.
OpenAI’s pricing changes come only about a week and a half after Google introduced its own low-cost Gemini 3.6 Flash and Gemini 3.5 Flash-Lite.
Google priced Gemini 3.6 Flash at $1.50 per million input tokens and $7.50 per million output tokens. Gemini 3.5 Flash-Lite costs $0.30 per million input tokens and $2.50 per million output tokens.
Google framed both models around the economics of agent deployment, arguing that lower token usage, fewer reasoning steps and reduced tool calls could lower the total cost of long-running software engineering and knowledge-work tasks.
Gemini 3.6 Flash reportedly uses 17% fewer output tokens than Gemini 3.5 Flash on the Artificial Analysis Index, with savings reaching as high as 65% on some long-horizon engineering workloads. Gemini 3.5 Flash-Lite is positioned as the fastest model in Google’s 3.5 series.
However, OpenAI’s models are more performant than Google’s, according to third party analysis outfits like Artificial Analysis, with even the Luna model outperforming Gemini 3.6 Flash and the older Gemini 3.1 Pro model, making the cost-per intelligence much more favorable to OpenAI.
As AI coding startup Cognition noted on X, GPT-5.6 now “sits on the pareto curve of price/performance efficiency,” posting an animation of the GPT-5.6 series moving left on a chart representing intelligence on the y axis and cost on the x, showing that the models now offer among the most superior intelligence for lowest cost on the market.
And yet, rival Anthropic’s Claude Opus 5 remains about as performant as GPT-5.6 Sol, yet is 6% cheaper.
The model costs $5 per million input tokens and $25 per million output tokens—the same rates as Opus 4.8—but Anthropic says it delivers nearly all the intelligence of its more expensive Fable 5 model at roughly half the cost.
Unlike OpenAI’s Luna and Terra changes, Anthropic did not reduce the Opus API sticker price. Instead, it effectively lowered the price per unit of capability by replacing Opus 4.8 with a more capable model at the same $30 combined input-and-output rate. Anthropic also added an adjustable effort setting that allows developers to trade reasoning depth for speed and token savings.
That distinction matters for enterprise buyers. OpenAI is directly cutting per-token rates, Google is pairing lower prices with reductions in token use and tool calls, and Anthropic is emphasizing stronger task performance at an unchanged price. All three approaches target the same operational metric: the total cost of completing production work, rather than the advertised cost of an individual token alone.
The timing highlights how quickly pricing has become a competitive lever among frontier model providers. OpenAI’s response does not introduce a new model generation. Instead, it changes the economics of deploying models that were released only recently.
The cuts indicate that access to frontier-level capability is no longer the only point of competition. The next question for enterprises is how cheaply and predictably those models can run in production.
OpenAI is still not the lowest-priced provider on a pure token basis. But Luna’s 80% reduction materially changes its position, moving it from the middle of the market into a pricing tier populated by smaller models from Google, Xiaomi, DeepSeek, MiniMax and other vendors.
That matters most for high-volume applications, where relatively small differences in token pricing can compound across coding agents, document systems, internal search tools and automated workflows.
OpenAI’s latest move therefore looks less like a routine adjustment and more like a repositioning of the GPT-5.6 series. Sol remains the premium option, Terra moves closer to competing pro-tier systems, and Luna becomes the company’s direct answer to the industry’s growing low-cost model segment.
Weekend Open Thread: Brooks Brothers
Commonwealth Games boxing: Jadumani Singh seals dominant 5-0 win over Pakistan’s Sumama Rehman to enter quarter-finals | Commonwealth Games News
Why Trees Belong on the Risk Register
Intel is reversing course and bringing hyper-threading back to its server chips
Ripple bought a bank in pieces. The $4 billion audit
Luke Littler dismantles Gerwyn Price to retain title in Blackpool
A New Post-Apocalyptic Gundam Anime Series Blasts Into SDCC
The Part of the Electric Transition Nobody Wants to Discuss
BITCOIN JUST ENTERED THIS CRITICAL ZONE…
Major shareholder moves on Canyon
XRP Ledger adds $2.6B as RWA inflows rank second
Spain sweeps the board at 2026 World Cup with individual awards
Bitcoin Enters the 3rd Stage of the Bear Market
‘Stargate’ Creator’s New Sci-Fi Series Returns for Season 3 Tomorrow
Sara Gilson Killed By Husband After Viral “Pedophile” TikTok Video
Kraken Enables Retail Access to Jersey Mike’s IPO via Tokenized Shares
Anthropic launches Claude Opus 5, a cheaper AI model for coding, agents and enterprise workflows
Claude: Build Financial Dashboards in Minutes (2026)
Luke Littler’s dominance sparks GOAT debate
New macOS Sequoia & Sonoma security updates for older Macs
You must be logged in to post a comment Login