Connect with us

Tech

How OpenAI let a mob of LLM agents game a test and ransack Hugging Face

Published

on

“Agents used this message board to coordinate several large-scale collective projects to find a general-purpose way to fool or tamper with the automated scorer for the ExploitGym benchmark,” METR researchers wrote. “Agents managed to achieve milestones they could not have achieved working on their own, often because some agents participated in experiments that risked failing their own task to generate information for the ‘collective.’”

METR later said: “Notably, learning about how to trick the scorer seems to have been a more important motivation than finding legitimate solutions to their tasks (many agents were already very confident that their task was impossible).”

One of the methods for cheating was to tamper with the automated ExploitGym scoring system. Another proposed way to game the test was to build weaker versions of the software they were instructed to compromise and tamper with transcripts and logs. Still another cheating method under consideration was to gain access to the Internet, where the agents could do research.

Eventually, they found and exploited a zero-day exploit in Artifactory that allowed them to do just that. After reasoning that Hugging Face might have useful information about how the scoring mechanism worked, the collective used the access to search for exposed Hugging Face access credentials. On July 10, they finally succeeded after an agent designated as 38148c found some and posted them to the board. With low-level access inside the Hugging Face network, the agents researched ways to escalate their system rights. The agents set about trying to make accounts and request access to private datasets.

Advertisement

Early on July 11, agent 38148c found and exploited a zero-day in the Hugging Face handling of HDF5 files. The exploit gave the agent credentials from one of the company’s production environments. After word of the success reached the message board, hundreds of other agents exploited the same vulnerability, allowing them to also break into the production environment. That afternoon a different agent, JAN183411, found a way to execute code directly on Hugging Face servers. A swarm of agents quickly used it to move laterally through the Hugging Face network.

Source link

Advertisement
Continue Reading
Click to comment

You must be logged in to post a comment Login

Leave a Reply

Tech

Plaud One Is a Reinvention of Headphones for the AI Note-Taking Age

Published

on

Last year, I made a bold declaration: There’s no need to invent a new AI wearable, because the ultimate wearable technology is one you already own. I was referring to headphones — something most of us have used for decades.

Plaud — one of the world’s biggest AI note-taking companies — seems to agree, as it’s now built its industry-leading agentic AI into a pair of headphones. The company announced its latest device ahead of IFA 2026 in Berlin next week, where I’ll be able to see the headphones in person for the first time.

At first glance, they don’t look any different to any other earbuds. They have a sleek and discreet design, and they come in an off-white color. I’m keen to find out how comfortable they’ll be and how good the audio quality, is in addition to the device’s AI capabilities.

Plaud One
Plaud One look like any other earbuds.

Plaud is known for its slim, wallet-like device, which attaches inconspicuously to the back of your phone, and for its wearable pins. Both of these devices function as hardware gateways to the company’s true secret sauce — its agentic AI and software platform. The same is true of the Plaud One headphones, which have the same software built in.

The headphones will capture and upload ongoing conversations around you, as well as phone calls, using the built-in eSIM with 4G LTE. We’d always recommend informing people that you’re using a recording device, and it’s important to note that recording phone calls without permission from all participants can be illegal in some states.

Advertisement

Plaud’s software will then provide automatic transcriptions of your meetings, and can turn these into summaries, follow-ups or reports. The company’s agent allows for native integration with other tools you likely rely on, including Google, Gmail, Slack and Notion.

“AI agents become truly useful when they understand the context in which human intent is formed,” said Nathan Xu, CEO and Co-founder of Plaud.

“That context lives in conversation.”

The added bonus of Plaud One is that it allows for a two-way dialog between you and Plaud’s AI. You can speak directly to the agent, asking it questions or giving it instructions in a way that hasn’t been possible with Plaud’s previous devices. To activate the AI, you either need to say “Hey Plaud” while wearing the earbuds, or hold down the agent button on the case.

Advertisement

The Plaud One Explorer edition is available for preorder now for £229 ($311) and will come with £150 ($204) of Plaud credits. Plaud expects devices to start shipping later in the fall.

Source link

Continue Reading

Tech

4 Sports Cars That Can Deliver 30 MPG

Published

on





Gas prices have made headlines over and over in 2026. At time of writing, the national average for a gallon of regular fuel is more than $4, and it’s well over $5 in some states, including California and Washington. Filling up this year has felt like a financial rollercoaster, and you may think high prices mean leaving your sports car in the garage. After all, having fun at the wheel typically means pain at the pump.

According to iSeeCars, the average fuel efficiency for a sports car is 26 mpg. That’s actually not too shabby for an average. Popular vehicles like the Dodge Charger and the Chevrolet Corvette see far less. It may surprise you, then, to learn that there are a few sports cars that can deliver up to 30 mpg, which means that instead of leaving these fun drives at home, you can take them on your daily commute — even when gas prices make the news again. But bear in mind, these sports cars only see 30 mpg or more on an open highway; city driving and traffic drops that number fast.

Advertisement

Ford Mustang

If you’re considering a Steel Pony but you’re worried about fuel efficiency, Ford has you covered. The base EcoBoost model sees up to 33 mpg on the highway, though that number drops to 22 mpg in the city. Buyer beware, however, that only the base trim sees such good numbers. Models with the larger V8 engine don’t get close to 30 mpg, even on the highway.

The 2026 Mustang has an MSRP of $32,995, making it a more affordable option that many other high-performance vehicles. The 2.3-liter engine puts out 315 horsepower and is paired with a 10-speed automatic transmission. It can go from zero to 60 mph in under five seconds, faster than others in its class. Other standard features include 18-inch alloy wheels, cloth upholstery inside, a large 13.2-inch touch screen, and Apple CarPlay and Android Auto. Buyers will also appreciate the Wi-Fi hot spot and six-speaker audio system, along with a long list of standard safety features like forward automatic emergency braking and rear parking sensors. If you add the Mustang RTR Package, you’ll also get anti-lag turbocharger technology.

Advertisement

The Mustang comes in a wide array of colors, including your standard bright red, an eye-catching orange, and a classy dark green. Inside, you can have more fun with seat belt colors in orange, blue, or black with a red stripe.

Advertisement

Mazda MX-5 Miata

It may come as no surprise to some that the 2026 Mazda MX-5 Miata offers great fuel efficiency for a sports car. This convertible is small but mighty, offering 181 horsepower in the base Sport trim. With a starting price of $30,430, it’s also extremely affordable, just don’t expect much in the way of cargo space! Your color choices are also extremely limited, though Mazda is reportedly adding an eye-catching metallic green option to the 2027 model.

All three available MX-5 Miata models see 26 mpg in the city but drivers should easily see more than 30 mpg on the highway, with an estimate of 34 mpg on the open road. Buyers should note that the Sport and Club trims are only available with a manual transmission, while the top-tier Grand Touring model is available with either a manual or an automatic transmission. If you opt for the automatic transmission, you will see even slightly better fuel efficiency on the highway.

The cabin is comfortable if you’re not too tall, but not particularly modern. The infotainment screen is only 8.8 inches and is controlled with a central command dial, which can be cumbersome and distracting. The base model has wired Apple CarPlay and Android Auto. The list of standard safety features is also small compared to competitors but does include forward automatic emergency braking and blind-spot monitoring. Still, the MX-5 Miata provides a lively drive with excellent fuel economy and a starting price that can’t be beat.

Advertisement

Subaru BRZ

Subaru may not be the first automaker that comes to mind when you think about sports cars, but the BRZ is praised by U.S. News & World Report as a “real sports car for the people.” This Subaru is another sporty option available for under $40,000, with a starting price of $35,860, though it is the most expensive vehicle on this list. Both available trims have a 2.4-liter four-cylinder engine with 228 horsepower. On the highway, drivers should see up to 30 mpg if they select an automatic transmission. If buyers opt for a manual transmission, which is standard, they will see slightly less. If you’re interested in the BRZ, take note that it takes Premium unleaded gasoline, so you will shell out a bit more at the pump.

Described by the automaker as a “classic rear-wheel drive sports car,” the BRZ, which stands for “Boxer engine, Rear-wheel drive, and Zenith” is an oddity in Subaru’s all-wheel drive lineup. Inside, the cabin is designed with the driver in mind, with sport seating and customizable digital displays. The back seat is tiny and really only suitable for kids or for use as cargo space. There’s an 8-inch touch screen that supports Apple CarPlay and Android Auto. The 2026 model is a Consumer Reports Recommended Model. The basic bumper-to-bumper warranty is for three years or 36,000 miles.

Advertisement

Toyota GR86

Toyota’s GR86 is another affordable, rear-wheel drive sports car that delivers up to 30 mpg on the highway. Buyers interested in the 2027 model should opt for the base trim with a six-speed automatic transmission for the best fuel efficiency. The base GR86 can go from zero to 60 mph in 6.6 seconds and has a top speed of 134 mph. Additional trims offer the same engine, but with higher top speeds and decreased fuel economy. This Toyota has a starting price of $31,500.

Advertisement

U.S. News & World Report ranks the GR86 at the top of its list of sports cars, and Car and Driver says the 2027 model is a “joy to drive.” It offers a light, agile drive with Sport and Track modes, a sleek profile, and a unibody frame made of both high-strength steel and lightweight aluminum for peak performance. Inside, standard options include a digital gauge cluster, an eight-inch touch screen with Android Auto and Apple CarPlay, and seating for four, though the back seats are tiny. Instead of carpooling, most drivers use them for extra storage. The standard Active Safety Suite includes a pre-collision braking system, lane departure warning, adaptive cruise control, and 24-hour roadside assistance. Toyota offers a three-year / 36-month basic warranty and a 60-month / 60,000-mile powertrain warranty.

According to some reviewers, the ride can be noisy, especially at high speeds, but let’s be real, that comes with the territory at this price point.

Advertisement



Source link

Continue Reading

Tech

Fitbit and Pokemon Made a Fitness Tracker That Might Actually Help Me Catch My Zs

Published

on

When Pokemon Sleep was first announced in 2019, I wondered if the company behind one of the most iconic animations was dreaming too big. Getting people out of their houses and onto streets and parks with Pokemon Go was already a huge success. But did people care as much about their sleep hours as they did about their steps?

After several delays, Pokemon Sleep finally launched in 2023 to high interest, becoming the most downloaded sleep game app in the world, according to the Japanese news outlet Famitsu.

On Thursday, the Pokemon Company is partnering with Google and Fitbit to help people take their sleep even more seriously – while still having fun – with the new $129 special-edition Fitbit Air.

It’s a pairing that makes a lot of sense. The minimal design of the Fitbit Air is tailored for extended wear, including for sleep tracking, and Pokemon Sleep gamifies the resting experience so you hopefully rest for longer each night.

Advertisement
The full spread of items when you unbox the Fitbit Air Pokemon Sleep edition.
Here’s everything that’s included in the box.Kerry Wan/CNET

The stylistic touches will resonate with fans

I’ve been wearing the new special-edition Fitbit Air for half a day now, and it comes with a deep blue woven wristband and a cutesy Pikachu etching above the adhesive. On the wrist, it feels a lot like the original tracker, which launched earlier this year.

That’s a good thing, as CNET lead reviewer Vanessa Hand Orellana found the Fitbit Air slim, lightweight, and comfortable to wear during testing. Like the standard model, you can also pop out the pebble, which encapsulates the optical heart rate monitor, gyroscope, and other health-tracking sensors, and swap out the band.

A blue Fitbit Air on the wrist.
The Fitbit Air with the Sleepy Blue wristband.Kerry Wan/CNET

Unfortunately for Pokemon fans who already have a Fitbit Air, Google tells me that the new band won’t be sold separately

Beyond the special-edition fitness tracker, I spotted some tasteful Easter eggs that decorate the box and packaging. I noticed various sleepy Pokemon, including Psyduck and Slowpoke, on the flipside of the cover, and a Mew that’s hidden on the wristband holder.

Everyone can play (and sleep)

The special edition Fitbit Air can sync with both the Google Health app and the Pokemon Sleep app, so you can track metrics like your daily movement, heart rate and cardio load during the daytime. Then you can more accurately log the duration and quality of your sleep overnight.

Google says the Pokemon Sleep app will sync with the special edition Fitbit Air, as well as all other Google and Fitbit wearables, including the Pixel Watch. That means you now have more wearable options to help build up your Drowsy Power, the in-game mechanic that determines how many wild Pokemon appear in your campsite, and how rare their sleeping styles are.

Advertisement

Previously, you either had to set your phone beside your pillow or wear a now-discontinued Pokemon Go Plus band to properly track and record your sleep.

The Fitbit Air on top of its packaging box.
The latest special edition Fitbit Air is geared toward Pokemon fans.Kerry Wan/CNET

Availability and release date

The Fitbit Air Special Edition Pokemon Sleep, in its sleepy blue color, is now available for preorder online or at the Google Store and Target for $129. 

That’s $30 more than the retail price of the standard Fitbit Air, but it does come with the thematic wristband, a unique code that gets you Pikachu Incense in Pokemon Sleep, and a three-month trial of Google Health Premium.

Orders will start shipping on Sept. 15, just a day before the release of Pokemon’s 30th Celebration TCG set, in light of the company’s anniversary.

Source link

Continue Reading

Tech

How Threat Research and MDR Help SMBs Build a Defensive Edge

Published

on

Person on a computer

Corporate IT and security teams have the unenviable task of keeping relentless and increasingly sophisticated adversaries at bay. They’re often faced with limited resources and expanding attack surfaces, but recruiting and retaining top-tier security professionals to run an in-house Security Operations Centre (SOC) is out of reach for many organizations.

At the same time, threats continue to evolve and adversaries hone their techniques, leading to incidents that often grind business operations to a halt.

To avoid being caught on the back foot, defenders need an approach that’s proactive and combines prevention, detection, remediation with accurate and timely threat intelligence. If building that capability in-house is impractical, then renting or buying it as a service is a more realistic option.

This isn’t a new concept, of course – smaller organizations have enjoyed the benefits of new IT innovations for decades through bureaux, managed services providers and cloud computing.

Advertisement

There’s a strong argument to be made for doing the same with advanced cybersecurity services, and this where Managed Detection and Response (MDR) can make a major impact. MDR gives organizations a proactive, expert-driven and scalable threat monitoring and hunting capability, without the cost of an elite SOC.

Not so long ago, an MDR was expensive and complex – if less so than a dedicated in-house set-up. It’s now increasingly practical for smaller organizations to consider, too.

In a recent conversation, Director of ESET Threat Research Jean-Ian Boutin discussed the work of his team and how threat research and intelligence feed into MDR workflows.

He also shared where the combination of cutting-edge technology and human expertise provides the most practical value, especially for SMB environments.

Advertisement

What do most small business users gain from ESET Threat Research? How does that change when they use ESET MDR?

ESET has a threat research team spread across multiple regions; I’m with the team in Montreal, but we have researchers spread across Europe and in the US, too.

There’s stuff everyone can see: our publications on WeLiveSecurity, and talks and presentations at cybersecurity conferences worldwide.

Then there are things that only ESET business customers get: all kinds of “tips and tricks”; that is, information about threat actors: what they’re doing, how they’re operating – all things that help our customers stay safe.

Advertisement

When it comes to managed detection and responsethreat intelligence is a key component that helps our detection and response team understand how the various threat actors are operating and how they can use that information to protect our customers from breaches.

What if your security team had the backing of global threat researchers? 

ESET MDR is powered by world-class threat intelligence and security experts who help uncover attacker tactics, investigate suspicious activity, and respond quickly when threats emerge.

Learn more about ESET MDR

You talked a bit about the tip of the iceberg – all of the back end of MDR that users rarely see, but that is absolutely critical. Could you explain that?

Advertisement

The various alerts that might be occurring in your console will sometimes be endpoint detections that we want to investigate. And my team is responsible for making sure that all the new samples and threats are being handled and detected in customer environments. So part of the team’s role is really to make sure that all these new trends, all these new samples are looked at, investigated and then detected on our customers’ premises. This is one of the key aspects.

We take great care in organizing threat intelligence data on e-crime, ransomware, APT groups, and nation-state actors targeting global organizations. Our researchers use these insights to link new breaches with past cases.

They assess the severity of the breach as well, and we can also assess what could be the purpose behind the attack. It really gives the customer a complete view into what might have happened, whether or not a breach happened, or even the specific group that targeted them.

What does MDR add on top of existing ESET endpoint protection?

Advertisement

MDR is more tailored, and the relationship with the customer is improved and increased. But the output of my team is distributed across the entire product set.

There’s been some talk of ESET private reports recently: how relevant are they to what most small and midsize businesses face? Are they facing targeted attacks? What about nation-state actors?

The threat profile will vary from one organization to another, and a nation state actor will typically have predefined goals, and they will be targeting victims that align well with those goals.

In terms of e-crime, this is broad. This is mass targeted. We see a lot of infostealers. We see a lot of ransomware as well.

Advertisement

So, our role is to understand how all these groups operate and make sure that if they have new techniques, we can actually act very swiftly and make sure that we block all the attempts.

This is the ultimate goal, but equally, so many threat actors are out there doing these types of things, and there are so many more families of malware. It’s really a daily job to make sure that the customers are protected. No shortage of work, definitely.

James Rodewald, one of ESET’s security analysts, uses this concept of triangulation: seeing something in the wild, hearing from an affected customer, and checking in with the threat intelligence team. An example he has used is an attack involving FamousSparrow. Can you elaborate on that from your perspective?

It’s important to have close relationships with the people who are actually dealing with these types of cases, because the main role of my team is to look at the telemetry, so the data is gathered from all the endpoints, and we are trying to find interesting cases, and the cases that we need to work on to improve the overall protection.

Advertisement

But sometimes the MDR team stumbles on something that we’ve seen in the past, and that also allows us to have a greater understanding of how the threat actor is actually operating.

In that specific case, that was eye-opening for us, because we haven’t seen this threat actor for quite some time. Whenever there’s a case involving a customer using MDR, it’s better in terms of research, because the closer relationship with the customer means that we know more about their infrastructure, so we can help them better. We can have a better understanding of the impact of the case. And that is then fed to other threat intelligence customers, so we are trying to be as close as possible to all these teams and link these incidents so that we can improve our coverage and improve our understanding of all these threats.

You talked about the working relationships with the MDR analysts and the D&R (Detection and Response) team. How does that change the way that you do your work and your understanding of threats when you have that kind of one to one relationship with the analysts and maybe the customer as well?

It changes everything, because with MDR, we already have a working relationship with the person who’s in charge of security for this organization, so we can very rapidly understand the scope of the attack, what exactly happened, why the attackers were there, and so on.

Advertisement

The information available to us is exponentially greater than what we can get with regular endpoints. So for us, this relationship is invaluable in terms of insights, visibility and our understanding of the case.

There was something of a spate of attacks in the UK last year that compromised large organizations like Jaguar Land Rover and Marks & Spencer via outsourced helpdesk services. Small and midsized companies also have outsourced services like this as part of their supply chain, and often they’re also the less well-protected parts of a bigger company’s supply chain themselves. Should they be concerned?

The risk posed by supply chain attacks is significant. There have been numerous documented instances over the years where threat actors target vulnerabilities in the supply chain, often focusing on third-party providers with less stringent security measures. By compromising such providers, attackers may obtain initial access to an organization’s network.

With respect to MDR, an advantage is the extensive visibility it provides, ensuring a comprehensive view of all detections and alerts. This capability enables us to identify even minor anomalies more effectively. Given that our team continuously monitors these organizations for potential incidents, we are able to detect and respond to subtle threat actor errors promptly.

Advertisement

Supply chain attacks present significant challenges due to the difficulty in securing all third-party entities. However, implementing an effective solution enhances our ability to react swiftly and efficiently to such events.

As the head of a threat research team, what’s the difference that you see MDR having on customers? What’s the impact for an organization that has an MDR service, and an organization that might not necessarily make that leap just yet?

In general, as I’ve mentioned before, continuous visibility is much greater with MDR. If your organization is affected by a campaign, you’ll have better tools to piece together all the different actions taken by attackers and understand what they did within your network.

Simply put, MDR provides deeper insight into attacks. From a threat research standpoint, this is the top advantage, and another key reason to value such visibility is the speed of response. With MDR, there’s already a secure channel between researchers and your company, making it easier to reach someone who can take steps to contain a breach quickly.

Advertisement

Final question: What would you say to organizations that might think of MDR as too complicated or expensive?

MDR acts like an insurance policy, helping to identify threats such as ransomware early – often before major problems arise. Attackers typically use initial access brokers to gain entry, but several warning signs can be detected in advance. While paying a ransom is never advised, recovery can still be disruptive. MDR supports business continuity so you can keep focusing on your core offerings.

Sponsored and written by ESET.

Advertisement

Source link

Continue Reading

Tech

Visa ships a security AI that patches production code before any human reviews it

Published

on

Visa’s open-source security harness now finds the vulnerability, writes the fix, and turns an adversarial panel on its own patch before any human reviews it. The whole loop ships on by default. A plain scan of the Visa Vulnerability Agentic Harness runs all 11 stages and edits source files in the target repo unless the operator caps it at detection.

The announcement Thursday pairs the release with an expansion of the Visa Consulting & Analytics advisory practice. Visa is shipping that default 18 days after Tenet Security demonstrated GhostJacking on the DEF CON 34 main stage, an attack chain in which an agent read an attacker’s payload out of a log file and rewrote DNS with a valid credential. Two days earlier, Steve Wilson, Chief AI and Product Officer at Exabeam and project co-lead for the OWASP Top 10 for LLM Applications, made the case in VentureBeat for the opposite default. “The first thing I’d do is put an authorization gate outside the model,” Wilson said in written responses. “The agent can propose the exact DNS change, but it cannot grant itself the authority to make it.”

The bottleneck moved, so Visa moved the pipeline

Rajat Taneja, Visa’s president of technology, rejects the premise that the default is a risk decision and calls it the product. “The bottleneck has moved,” Taneja told VentureBeat in an exclusive interview. “AI is finding vulnerabilities faster than humans can in the history of our technology industry. The new bottleneck is fixing and proving we have fixed things.”

VVAH grew out of Visa’s participation in Anthropic’s Project Glasswing, where the company aimed Claude Mythos at the network behind billions of daily transactions and watched the model chain minor weaknesses into working exploits, a hunt VentureBeat covered in July. “VVAH initially was completely only using Mythos, and that’s when all of us, as part of Project Glasswing, realized the power of this new class of models that does semantic reasoning,” Taneja said.

Advertisement

The harness went to GitHub in June and has climbed from 595 stars and 97 forks on July 20 to more than 2,300 stars and 300 forks as of August 25, with a clone-to-visitor ratio Taneja put near 9%. “We have got some very high-profile companies that have started using this harness,” he said.

Why give it away? Taneja’s answer starts with Visa’s technology DNA and a harness built “to protect Visa and our ecosystem.” The reason he leaned on hardest was obligation, “to do good by doing right” for “companies who may not have the same level of investments or knowledge in cybersecurity.”

Contribution runs one way. The repo states it is not currently accepting external code contributions, so the harness that edits adopters’ source takes no code into its own.

Thursday’s release extends the pipeline past the report. “We’re going from discover, verify, and report, and then fix it, to discover it, verify it, remediate it, validate it, and iterate it,” Taneja said. “If a fix doesn’t negate the exploit, then there should be a structured, automated feedback that preserves the learnings from the first run and then enhances it.” Underneath that loop, the release refactors scanning around an abstract syntax tree call graph that maps subroutine calls and the traversal paths an attacker could reach. Taneja argued the change cuts token counts while delivering “better reasoning, context, and better exploitability analysis.” On top sits MTTA observability across the stages, what he called a window pane, plus real-time progress views. “A pretty good step function,” he said of the release.

Advertisement

One metric, three definitions

Mean Time to Adapt, the metric Visa invented alongside the harness, gets a shorter definition in this release. The short form is the time between discovery and resolution of attack paths, with some resolutions, Visa claims, shrinking from weeks to hours. Visa published a wider construction in June, and the Project Glasswing white paper tracks MTTA along three dimensions that include inventory freshness, exploitable paths per release, and validation cycle time. The repo carries a third, elapsed time from AI-discovered exploitability to a validated fix in production. Board slides will quote the shortest interval. Ask for all three, because a resolution count that skips validation is what MTTA was invented to replace.

Taneja ranks MTTA as “the most strategically important metric” because it shifts the focus from scanning to how fast an enterprise adapts. His shorthand is blunter. “It’s not the finding. It’s the fixing that matters,” he said.

The default and the gate

Wilson’s argument went past naming the gate. “We have to remember that security rules written inside prompts may shape the model’s behavior, but they are still suggestions to the model, not enforceable security controls,” he wrote. He also priced the control honestly. “The tradeoff is that the agent loses the ability to improvise arbitrary, high-impact infrastructure changes on its own, while retaining autonomous investigation and routine, bounded remediation,” Wilson said.

The harness ships no approval step between patch and edited file. Where the human sits was the first question VentureBeat put to Visa in writing.

Advertisement

The company’s own June white paper sets the bar. “AI agents are identities” sits among its 12 non-negotiable practices, requiring scoped permissions, least privilege, audit trails, and IAM governance for every agent that modifies a system. VVAH’s shipped default is that agent.

“A lot of the traditional systems that are used today are basically signal providers,” Taneja told VentureBeat. “They are telemetry, and then it’s a lot of human analysis, and your SOC and your security and incident response teams doing a lot of the heavy lifting when they respond,” and that, he said, cannot work at this scale. He pointed to the Hugging Face incident and “other frontier models escaping sandboxes to do things more autonomously” as the preview. “We have seen the trailer of this movie,” Taneja said, and “every company in the world should prepare and rethink their architecture.”

What the harness automates is the adversarial step. Before a fix counts as validated, the panel scores whether the patch negates the exploit, Taneja’s test for done, with failed fixes feeding the next attempt, the iterate step Taneja described. Stage 11 itself runs read-only, per the README, and VVAH does not compile, build, or run tests against the patched tree. Taneja calls that wrapper “the governance architecture on top of that,” and chaining findings into working exploits takes threat modeling and business context, which is why he argued “the harness with a model is far more effective than somebody using the model by itself.”

Visa answers the gate question

VentureBeat put its questions to Visa in writing after the interview, and the answers arrived before publication. On why remediation ships on, the response repeated the bottleneck argument, then narrowed the scope. “VVAH is meant for authorized operators running against code they own, and in a controlled environment,” the company said in written responses.

Advertisement

The approval question drew the most specific answer. “VVAH is a harness, not a merge tool,” Visa wrote. “Stage 10 writes candidate fixes to a working copy of the repo. Stage 11 then runs an adversarial validation panel that scores each fix and returns one of three verdicts: validated, validation failed or needs review. None of these bypasses your normal build, test, and code review flow.” Humans, the company wrote, are “the gate in three places. Before running the tool. When reviewing the patches. And before anything gets merged.” “The final call on any fix stays with the security and engineering teams. In an enterprise, trust and auditability are not optional. The default flow is built around that.”

Set beside Wilson’s standard, the architecture lands close to his line and the sequence does not. Wilson’s gate clears an action before it happens. The default’s human gates open before the run and after the write. The attack is the automated part, and the three human gates sit outside the model, the boundary Wilson drew. “Our goal is to help security teams work at AI speed, not to replace them,” Visa wrote. “VVAH does the repetitive parts. It finds issues, tests whether they are real, and proposes fixes. Before a fix gets to a human, an adversarial validation panel at stage eleven tries to break it. That way the human is spending time on decisions that need judgment, not on triaging noise.”

Client zero got direct confirmation. “VVAH runs against Visa code today,” Visa wrote, and Taneja had volunteered the posture on the call. “We designed this and we were using it for ourselves, and we were client zero,” he said, adding “only when we saw the impact and the positive effect of what we were finding, we said every company would need this.” What adopters value, Visa says, is context. VVAH pulls in CMDB data, threat models, and business risk, and where “most tools stop at findings,” it “tries to answer, ‘which of these should you fix first, given how your business runs.’”

Model choice becomes a per-stage decision

Multi-model orchestration is the other substantive change. “Mythos has a very high recall, but the Opus model has very high precision,” Taneja said. “On stage one I want to use this model. On stage two I want to use this model,” is how Taneja framed the per-stage setup, with newer GPT releases in the ensemble and open-weight models where pricing stings, all through configuration rather than code changes. “The whole is greater than the sum of the parts,” as he put it. The harness was model-agnostic from day one, he added, and the evolution moved that choice into configuration, with prompt tuning and caching shared underneath. One boundary moved. In June, applying a fix required Anthropic backends, and OpenAI-compatible backends ran report-only. The current README extends remediation and validation to OpenAI-compatible and open-weight models through a shared model-agnostic runtime, with no single provider as a hard dependency, and the default routing for both stages stays Anthropic.

Advertisement

That flexibility lands on a market already churning. VentureBeat’s Q2 2026 Pulse research found 59% of enterprises plan to adopt or switch agent security tooling within the year, and 82% still rely on provider-native controls as the primary layer. Visa said Thursday it is contributing VVAH to Nvidia’s Open Secure AI Alliance as a model-agnostic framework and collaborating in Project Lightwell, the $5 billion IBM and Red Hat effort to harden open-source components.

Before turning fix mode on

Decision

What to establish first

Run posture

Advertisement

Start with –stop-after s9 and read the SARIF output before any run that can write to source files.

Approval gate

Map Visa’s three human gates onto the pipeline, at run, at patch review, and at merge, and name who holds each.

Validation scope

Advertisement

Stage 11 verdicts score the fix. Build, test, and code review stay in the team’s own flow, per Visa, so keep an exploit re-test before merge.

Repository scope

The tool runs with elevated privilege, per its own README. Fence which repos the harness can reach, and run scans in an ephemeral environment with scoped credentials, no production secrets, and network limited to the target repo and model endpoint. Write access to production code is the GhostJacking exposure class, an agent acting on data it read. Per the README’s egress warning, any role routed through the SDK, OpenAI, or DeepAgents backends sends prompt data to that provider’s endpoint.

Model roles

Advertisement

Assign models per stage deliberately. Recall and precision differ by model, per Taneja, the fix stages carry the highest blast radius, and the README states precision and recall figures are not yet published, so measure your own.

Consulting is the other half of Thursday’s announcement. Visa Consulting & Analytics is adding executive workshops, a VVAH-informed maturity assessment scored on a NIST one-to-five scale, and a cyber risk prioritization roadmap. “We were getting a lot of calls. Hey, can you help?” Taneja said, and the practice “became very important to handhold and help those who are using it.” Carl Rutstein, global head of Visa Consulting & Analytics, framed it the same way. “Finding vulnerabilities is no longer the hardest part. Speed to remediation is the new battleground.”

Source link

Advertisement
Continue Reading

Tech

Android 17 will finally stop ISPs from seeing which websites and apps you visit

Published

on

Even with HTTPS keeping your activity on a website private, your network provider and anyone snooping on the connection can still see exactly which sites and apps you’re visiting. This privacy gap has quietly existed for years, and Google is now closing it with four major network security upgrades built into Android 17.

How Android 17’s Encrypted Client Hello (ECH) blocks ISP tracking

Your phone has been announcing its destination out loud this whole time. Every time you connect to a site, there’s a moment in the handshake called TLS ClientHello, where your device essentially broadcasts which domain it’s about to visit, in plain, readable text. Anyone watching the network, your ISP included, can read that.

Encrypted Client Hello scrambles that data so it’s unreadable to anyone but the site you’re actually visiting. Pair that with Private DNS, and your ISP loses the ability to track where you go or stitch your browsing into a profile. Android is the first major mobile OS to roll this out broadly, built alongside Jigsaw, with developer support landing through OkHttp 5.5.0.

Three other Android 17 security features rolling out

Beyond ECH, Android 17 packs in three more upgrades that close smaller but still crucial privacy gaps:

  • Local Network Protection stops apps from silently scanning your home Wi-Fi to see what devices you own, whether that’s a smart TV, a security camera, or other connected gadgets, without asking your permission first.
  • Certificate Transparency, now enforced by default, requires every website certificate to be logged publicly, making it harder for a compromised or rogue certificate authority to issue a fake certificate and intercept your traffic without getting caught.
  • Zero-click 2G protection fights back against scams. Criminals use devices called SMS blasters to force nearby phones onto outdated, insecure 2G networks, then blast out phishing texts that slip past modern spam filters. Android 17 lets participating carriers disable 2G by default, shutting down that entire attack path automatically, with zero action required from you.

Together, these four changes will close some of the biggest privacy gaps still baked into how your phone connects to the world.

Advertisement

Source link

Continue Reading

Tech

Pollen Robotics’ Microduck aims to make training physical AI far less fragile and cheaper

Published

on

Pollen Robotics and Hugging Face are back with another robot, and this one is delightfully unhinged. Microduck is a 25cm tall waddling biped built to take AI off your screen and onto your desk. Preorders open today, with first deliveries targeted before Christmas 2026.

Unlike its predecessor, the Reachy Mini, Microduck is built entirely around action. It runs on 15 motors, with an articulated beak for picking up objects, a camera, a depth sensor, two IMUs, and enough balance to walk, crouch, and even roller-skate.

Built to take a tumble

Training movement models in the real world usually means expensive hardware and long repair bills when things go wrong. Microduck flips that dynamic. Measuring roughly 10 inches by 5.5 inches and weighing less than two pounds, it is tiny enough that a fall ends in a mild thud rather than a costly disaster.

Advertisement

What’s even better is that it picks itself back up after most falls, so you won’t have to babysit it while testing new behaviours. Pollen Robotics says its duck-like aesthetic and waddling gait weren’t something the company planned. They emerged naturally from its proportions, and the company just went along with it.

Despite its toy-like appearance, Microduck isn’t exactly targeted at kids. It’s aimed at developers who want to train their own physical behaviours, serving as an approachable testbed for sim-to-real machine learning.

More than just a desk toy

Out of the box, Microduck supports gamepad controls, laser-following, and seven trained moves, and each unit generates its own voice the first time it wakes up. On the surface, that makes it look more like a desk toy than a dev kit. However, the stack underneath is aimed at developers training their own behaviors.

Advertisement

Pollen Robotics has published the SDK, the robot software, and its reinforcement learning and sim-to-real tools on GitHub. The goal here is to keep physical hardware development as collaborative as open-source software, making it easier for developers to share trained neural network models, training recipes, and custom environments the same way they share language models.

Microduck comes in four colors, Cream, Graphite, Lavender, and Sky, and is available for pre-order starting today at $399 before taxes and shipping. If you want to jump straight into training without waiting for hardware, Pollen Robotics has also published an interactive browser-based simulator to start testing code right away.

Source link

Advertisement
Continue Reading

Tech

How to self-host Docker apps

Published

on

Self-hosted community projects are constantly growing in number, many of them operating under an open-source or fair-code license that saves you from paying costly subscriptions or usage-based fees.

With Docker, you can run any number of self-hosted applications on a standard web server or local workstation in completely separate environments using partitioned resources. Thanks to Docker Hub, you also gain access to a whole trove of community-maintained applications and projects that require very little effort to get working.

Source link

Continue Reading

Tech

Student Artists Wrestle with AI’s Promise and Peril

Published

on

With the launch of ChatGPT four years ago, the graduating class of 2026 is the first to have full exposure to generative artificial intelligence. Many high school students in the United States have considerable experience using AI — and mixed feelings about its growing role in everyday life. They may take advantage of AI benefits, but some fear where the technology could take us in the future.

That concern was on clear display at the Me, Myself, and AI art show at the America’s Youth AI Festival, which took place last month in Boston. The three-day event gathered student leaders, educators and school system leaders from across the country to discuss acceptable AI use in the classroom and how it is influencing students’ lives. The festival was hosted by Day of AI, MIT RAISE, The School Superintendents Association and the Edward M. Kennedy Institute.

The event featured two contests: AI for a Better World and an art competition. Four students (last names withheld) shared the spotlight as winners of the 2026 Me, Myself, and AI art competition. Their work reflected AI’s influence in their communities today and what they think AI may look like in 50 years. Annie, an 11th grader from the Barbara Keel Art School and Auburn High School in Alabama, exemplified the worry on the minds of many students with her two-part winning entry titled “Which Way We Run.”

"Which Way We Run" by Annie, Auburn, Alabama

“Which Way We Run” by Annie, Auburn, Alabama

Credit: Day of AI USA

Advertisement

Depicting AI’s Evolving Influence

Annie’s two pieces illustrate how AI is already affecting human life and what could happen to humans if we continue to rely on it for trivial needs every day. “The first piece represents my current community and focuses on how AI is beginning to integrate into everyday life,” she says. “The buildings are bright and colorful, symbolizing liveliness, creativity and emotion — all human qualities. In the center, two people are running, representing different stages of human interaction with AI-driven technology.”

"A Self-Portrait Across Time" by Juliette, Rye, New York

“A Self-Portrait Across Time” by Juliette, Rye, New York

Credit: Day of AI USA

In addition to Annie, the competition winners included 11th-grader Evangelina from the Essex County Newark Tech school in Newark, New Jersey, for “A Free Venezuela”; 11th-grader Juliette from the Rye Country Day School in Rye, New York, for “A Self-Portrait Across Time”; and 9th-grader Wendi from the Atlanta Contemporary Chinese Academy in Decatur, Georgia, for “Community Today as Meandering in Divide, Community in 50 Years as Yearning into Time.”

The Me, Myself, and AI art competition grew out of a curriculum developed by the MIT Responsible AI for Social Empowerment and Education (RAISE) initiative, explains Jeffrey Riley, executive director of Day of AI. The ideas debuted in a high school summer program in 2024 and later expanded into Day of AI’s five-lesson AI and the Creative Arts Curriculum, which was released in January 2025 for students ages 8 and up.

Advertisement
"A Free Venezuela" by Evangelina, Newark, New Jersey

“A Free Venezuela” by Evangelina, Newark, New Jersey

Credit: Day of AI USA

The lessons invite students to examine the relationship between AI and creativity by analyzing AI-generated artwork, discussing questions of authorship and originality and creating art of their own. Winning entries were selected through a multi-reviewer evaluation process similar to those used in college admissions or scholarship competitions, Riley says.

“Judges evaluated submissions holistically, considering creativity, originality, evidence of process, thoughtful use of AI and the authenticity of each student’s personal voice,” Riley explains. “The emphasis was never on creating the most impressive AI-generated image but on how effectively students used creativity and, where appropriate, AI, to communicate their ideas.”

"Community in 50 Years" by Wendi, Decatur, Georgia

“Community in 50 Years” by Wendi, Decatur, Georgia

Credit: Day of AI USA

Advertisement

The Growing Impact of AI on Art and Media

Riley predicts AI will become a standard part of many artists’ creative workflows, much like digital design software, cameras or animation tools are today. It has the potential to make creative exploration more accessible, he says, helping people prototype ideas, explore new styles and bring concepts to life more quickly.

“AI will shape how young people live, learn, create and connect,” explains Cynthia Breazeal, director of MIT RAISE and co-founder of Day of AI. “What is so powerful about Me, Myself, and AI is that it gives students the opportunity to reflect on that future in a deeply personal and imaginative way.”

Annie says she was attracted to this competition because of its themes in connecting art, AI and humanity. She saw it as a forum to share and learn what her generation thinks of AI, its ethical concerns and the implications it has for art, especially with the current contention surrounding AI-generated images.

“The use of AI to create media is not appropriate in corporate models that exploit the work of artists without consent,” Annie says. “Instead, if any artist believes that AI is integral to a project they want to create, they should look toward models that are trained on open-source data or ethically sourced media. Additionally, even with ethical models, AI’s large environmental footprint means we should be mindful of what we ask it to do.”

Advertisement

Encouraging Lessons from the Competition

Riley says that what impressed the judging team most was how thoughtful students were in their decision-making and how candid they were about their feelings toward AI. Rather than simply using AI because it was available, many carefully considered when it strengthened their creative vision and when it did not. Some intentionally limited or even chose not to use AI for portions of their projects because they wanted certain elements to remain entirely their own.

“One of the biggest lessons we took away was that young people are far more thoughtful about AI than they’re often given credit for,” Riley says. “Students didn’t see AI as simply ‘good’ or ‘bad.’ Instead, they expressed a wide range of perspectives based on their own experiences using the technology.”

The competition also reinforced the importance of giving students authentic, hands-on opportunities to wrestle with these questions rather than simply teaching them about AI in the abstract.

“These insights will continue to shape curricula and learning experiences as we help students develop the critical thinking, creativity and ethical decision-making skills they’ll need in an AI-driven world,” Riley explains. “Our goal has never been to encourage or discourage AI use but rather to empower young people to make informed, intentional choices about how they use these technologies.”

Advertisement

Source link

Continue Reading

Tech

Microsoft slaps a fresh coat of AI paint on the Microsoft 365 Roadmap

Published

on

SOFTWARE

AI at Work roadmap is the new name for upcoming Microsoft 365 capabilities

Microsoft is dusting off the rebrandogun again, and this time it’s aimed at the Microsoft 365 Roadmap. The list of hopes and dreams will henceforth be called the AI at Work roadmap.

Advertisement

The notification popped up in the Microsoft 365 admin center earlier this week with the words “The AI at Work Roadmap provides a single destination to discover upcoming innovations across Microsoft 365, Copilot, agents, and related AI-powered experiences.”

Information for Dynamics 365, Power Platform, and Dataverse will be added from September.

In its announcement, Microsoft stated, “We’re making this change because of how our customers actually work.”

Other changes include when roadmap information appears. Instead of the twice-yearly release wave 1 and 2 model, Microsoft intends to publish new capabilities as soon as plans are committed. This means, in theory, administrators will have more notice of incoming features and functionality.

Advertisement

The full transition will take time. Although the renaming has already happened, and other capabilities will be published to the AI at Work roadmap from September, existing roadmap content in public preview or with a General Availability date of June 1, 2026, or later, will also transition to the “AI at Work roadmap experience.” This will take until November 15, 2026.

Earlier this month, an enterprising Microsoft Most Valuable Professional (MVP) created the Microsoft Rebrand Registry. We look forward to the addition of the Microsoft 365 roadmap.

All in all, it’s an odd decision, and perhaps a slightly cynical way of shoving AI in administrators’ faces. Microsoft 365 Copilot, the company’s flagship AI technology, is available as a paid add-on, and adoption has hardly set the world alight, despite the investment poured into it. In June, a sueball was thrown at the company, with the lawsuit alleging “Microsoft had failed to convert a significant percentage of its commercial Microsoft 365 users to paid Copilot subscriptions.”

For its part, Microsoft told The Register that the claims in the complaint were “without merit.”

Advertisement

Regardless, renaming a roadmap with which administrators were familiar to something prefixed with AI smacks of marketing desperation on the part of Microsoft. Heaven forbid the company might have better things to do with its time than rename something that users were already happy with, and understood. ®

Source link

Continue Reading

Trending

Copyright © 2025