Connect with us

Tech

Wizards of the Coast reveals first look at ‘Warlock,’ the next video game based on ‘Dungeons & Dragons’

Published

on

(Wizards of the Coast/Invoke Studios press image)

Wizards of the Coast and its subsidiary Invoke Studios released a first-look gameplay trailer for Warlock: Dungeons & Dragons, a new open-world action game based upon the systems and setting of Wizards’ popular tabletop role-playing game.

Wizards, based in Renton, Wash., is a subsidiary of Hasbro and the company behind the D&D and Magic: The Gathering franchises. It has become one of the toy giant’s most reliable profit engines. Hasbro and Wizards are now looking to video games for additional growth, with mixed results; Wizards’ recent video game efforts have included both the generational success of Baldur’s Gate 3 and significant flops such as 2021’s Dark Alliance.

Warlock‘s first trailer was released during Opening Night Live at the Gamescom conference in Cologne, Germany, which has become one of the hallmarks of the video game industry’s calendar, particularly in Europe. Clearly, someone at Wizards and/or Hasbro has a great deal of faith in Warlock.

Warlock is set in the brand-new fantasy world of Darkon. Players take the role of Kaatri (Tricia Helfer, Battlestar Galactica), a combat veteran and swordswoman who makes a pact with a magical entity to attain magical powers. Specifically, Kaatri’s deal is with Tasha (Maggie Robertson, Resident Evil Village), an archmage from the world of Toril and one of D&D’s longest-running signature characters.

In exchange for whatever Tasha wants, which is not yet revealed, Kaatri becomes a warlock: a spellcaster who derives her magic from her link with Tasha, rather than the other available sources in D&D’s cosmology. Kaatri’s newfound powers include signature D&D abilities such as telekinesis, short-distance teleportation, and detecting sources of magic.

Advertisement

As such, the player can approach every situation in a number of ways depending upon their opportunities and preferences, from open conflict to stealth to clever manipulation of their environment. You can go in swinging with Kaatri’s sword, telekinetically shove enemies off of cliffs, or take an opponent out at long range with D&D warlocks’ signature eldritch blasts.

Invoke Studios, based in Montreal, was founded in 2022 as a reorganization of the team behind the 2021 co-op action-RPG Dungeons & Dragons: Dark Alliance. Warlock is its debut project under the Invoke Studios brand.

To make it to market, Warlock and Invoke have survived two leadership shakeups at Wizards and several waves of layoffs and shutdowns, including the cancellation of two other D&D-based video game projects in early 2023.

Warlock: Dungeons & Dragons is scheduled for release in 2027 for PC, PlayStation 5, and Xbox Series X|S.

Advertisement

Source link

Continue Reading
Click to comment

You must be logged in to post a comment Login

Leave a Reply

Tech

M6 Mac mini vs M4 Mac mini: Compact powerhouses compared

Published

on

Apple has refreshed the Mac mini adding the M6, in an unusually timed update. Here’s how the new model specs compare against the M4 version it replaces.

Two small silver desktop computer boxes side by side, front view, with USB-C ports and power indicators, against a pastel gradient teal and pink background
M6 Mac mini vs M4 Mac mini

In a rare move, Apple has updated the Mac mini just before its main iPhone event, rather than during its usual fall Mac window. The big change is the shift from the M4 generation it has been on for almost two years.
That shift isn’t to the M5, but to M6.
Continue Reading on AppleInsider | Discuss on our Forums

Source link

Continue Reading

Tech

Air Draws the Picture, a Soft 64-Pixel Screen That Runs on Nothing but Vacuum

Published

on

Air-Powered 64 Pixel Screen
Soft silicone dots rise and fall across an eight-by-eight grid, forming shapes and simple patterns without a single glowing diode or liquid crystal. Soiboi Soft, the maker behind the project, built this entire display from air pressure differences, 3D-printed plastic, and thin membranes that act like valves and memory cells. Atmospheric pressure stands in for a binary zero. A vacuum stands in for a binary one. Once a pixel settles into place it stays there on its own, no continuous power required.



Soiboi Soft began with some small experiments, including a four-by-four model and a seven-segment clock capable of pneumatically storing its state. Scaling up to sixty-four pixels, however, necessitated significantly tighter tolerances and a more sophisticated control system. It turned out that you needed sixteen small solenoid valves to feed the eight row and eight column lines. The fluidic logic integrated into the stack determined which pixels turned on and off.

Sale


Dragon Touch 15.6″ digital calendar chore chart – 1080P full HD interactive touchscreen, smart family…
  • 【All-In-One Smart Family Calendar】: Dragon Touch digital frame effortlessly organizes and tracks every family schedule with the crystal-clear…
  • 【Easy Setup and Auto-Sync】: Enjoy a user-friendly design that our smart picture frame allows for quick setup—just plug in, connect to Wi-Fi, and…
  • 【Interactive Chores Chart & Dinner Planner】: Our digital calendar keeps housework organized and motivates family members, especially children, to…

Air-Powered 64 Pixel Screen
Things work like this: you hold a row and a column under vacuum pressure at the same time, and a local silicone switch opens, emptying the pixel chamber. The flexible dot on the surface pulls inward, resulting in a visible, touchable depression, similar to a pit. Close the switch, and the vacuum is trapped like a capacitor that never needs to be recharged.

Air-Powered 64 Pixel Screen
Layers should be stacked properly, but not too perfectly, because being too perfect is ineffective. You received a translucent PLA baseboard that houses the control chambers. The logic membrane is made up of a special silicone sheet, which moves to open or close the passageways as needed. On top of it, there’s a pressure routing board, followed by a final piece of molded silicone with the sixty-four raised dots that make up the visible surface.

Air-Powered 64 Pixel Screen
Everything is transparent, and you can see the internal channels in action. Getting it to work was a pain because consumer FDM printers are terrible at producing airtight parts, so he had to slow down the printing, raise the temperature, over-extrude some, and even put some concentric ridges around each chamber to help concentrate the sealing force. Then he’d vacuum clamp the boards in TPU sleeves and anneal them to keep them from warping, resulting in a soft, slightly yielding surface that you could actually touch, as it pushed back against your fingers.

Air-Powered 64 Pixel Screen
The images are built up row by row. The controller, which may be a gamepad or something, sets the column pressures and then activates each row one at a time. All of the pixels read and lock in their values. So simple shapes, numbers, letters, and even a basic game of Snake arise. Refresh rates aren’t super high, and the resolution isn’t great compared to any modern screen, but that’s all you’d expect. Once the valves have done their thing, the display doesn’t use any electricity, and the same vacuum network that writes the image also gives the surface its living quality, as the dots move, hold, and respond to touch in a way that feels exactly like pixels should.
[Source]

Advertisement

Source link

Continue Reading

Tech

Xbox layoffs fallout: ‘South of Midnight’ creator Compulsion Games successfully goes independent

Published

on

(Compulsion Games image)

One of the video game studios impacted by Xbox’s layoffs in July has successfully reclaimed its independence, as well as control over its intellectual property.

Compulsion Games, headquartered in Montreal, was founded as an independent studio in 2009 and acquired by Xbox in 2018. Its one release as a member of the Xbox Games Studio network was 2025’s South of Midnight, an action/adventure game set in a magical Deep South.

In July, Microsoft announced the first wave of a planned 3,200 job cuts throughout its Xbox department, alongside plans to spin out or shut down five of its studios. Compulsion Games was one of those five, alongside Double Fine Productions (Psychonauts), Ninja Theory (Hellblade), Undead Labs (State of Decay), and Arkane Studios (Deathloop, Dishonored).

Subsequently, on Aug. 20, Compulsion CEO Guillaume Provost revealed in an interview with GamesBeat’s Dean Takahashi that Compulsion’s management had successfully reacquired the studio, its staff, and the South of Midnight IP on Aug. 11.

South of Midnight is still available via its previous storefronts, including Steam and the PlayStation Network, but is currently self-published by Compulsion.

Advertisement

Provost told GamesBeat that no layoffs had been made at Compulsion as it transitioned to independence, and most of the team elected to stay together.

As for the other studios affected by Xbox’s July 6 layoffs:

  • Double Fine Productions, headquartered in San Francisco, confirmed on July 28 that it had laid off 23 employees to return the studio to a “sustainable size.” It is once again fully independent and has control of its IP, such as Psychonauts, and will be exhibiting in Seattle on Labor Day weekend as part of the Penny Arcade Expo.
  • Ninja Theory, from Cambridge, England, was one of the more unexpected shutdowns, as it had debuted Senua, a third entry in its Hellblade series, only a few days before the layoffs announcement. It has reportedly been spun off from Microsoft and will continue work on Senua under an unspecified new owner.
  • Likewise, Seattle’s Undead Labs is currently under unidentified new ownership and still plans to release the long-anticipated third entry in its zombie survival series State of Decay at some point next year.
  • Finally, Arkane’s status has yet to be firmly established. It formerly consisted of two studios, in Austin, Texas and Lyon, France, but its Austin office was closed down as part of a wholly separate wave of Xbox layoffs in May 2024. Several of the affected employees in Texas, including former studio head Harvey Smith, announced on Aug. 19 that they’d founded a new company, Black Pony Immersive, with plans to create new games in the same “immersive sim” subgenre as Dishonored.

Xbox is currently exhibiting at the Gamescom conference in Cologne, Germany.

Source link

Advertisement
Continue Reading

Tech

‘The right approach is for people to deeply control the future’: OpenAI CEO Sam Altman is worried about AI being controlled by just a few powerful players

Published

on


  • OpenAI CEO Sam Altman is worried that powerful companies are dominating the evolution of AI
  • Altman has also expressed concern over AI’s potential for acting autonomously and possibly escaping human control
  • The comments are believed to be a reference to his rival, Anthropic CEO Dario Amodei

Are too few companies wielding undue influence on the direction of AI? That seems to be the belief of OpenAI’s Sam Altman, who has warned against a scenario where the people don’t have a say in how AI changes the world. Speaking to the David Senra podcast, Altman expressed a desire to see “society and the models… co-evolve.”

While this may seem desirable, a byproduct of this thinking might be harder to accommodate. Altman also expanded on his view that governments should not regulate AI, expressing a view that prefers the existence of liberty over the possibility of safety.

Source link

Continue Reading

Tech

Lab Supply Companies Have Been Selling Antibodies Using Manipulated Images

Published

on

An anonymous reader quotes a report from Ars Technica: Antibodies play a central role in your body’s immune defense. But they’re also nearly ubiquitous in biological research, being essential for several widely used lab techniques. While some researchers need to go through the process of producing their own antibodies, antibodies to many key proteins are commercially available and can be ordered for overnight delivery. In May, however, Reese Richardson, a post-doc at Northwestern University, discovered that some of the images used to demonstrate how these antibodies performed in lab experiments had been subject to image manipulation — the sorts of changes that would get a paper retracted if they had appeared in the academic literature. The manipulations ranged from removing background noise to copying and pasting data to fabricate results. Now, he has done a more exhaustive search and found problematic manipulations in images used to market over 17,000 commercial antibodies.

[…] The consequences of this discovery are pretty significant. If you buy an antibody believing it will generate clean, clear data and find it doesn’t, many researchers will assume the problem is them. That often leads them to spend time troubleshooting the procedure and trying slightly different conditions to get useful data. If the antibody can’t produce those results, all the troubleshooting will have been a waste of time and money. Nature’s coverage of the new findings includes quotes from several companies that sell these antibodies that, paraphrased, range from “we’ll look into it” to “we don’t think it’s a problem.” Researchers who have spent time frustrated while failing to get an antibody-based experiment to work may think otherwise.

Read more of this story at Slashdot.

Advertisement

Source link

Continue Reading

Tech

Instagram’s new editing tool ‘First Draft’ turns your clips into a Reel in seconds

Published

on

Staring at a pile of unedited clips is often the exact moment we give up on posting entirely. Instagram just tackled that problem head-on, by rolling out First Draft, a new one-tap tool that assembles your raw footage into a finished Reel in a few seconds. For now, First Draft is rolling out exclusively on iOS, with no word yet on when Android users might get access.

How does First Draft create Reels for you on Instagram?

The process starts just like making any other Reel. You select clips from your camera roll, but instead of manually trimming everything yourself, you tap the new First Draft button instead. You can also access it directly from Instagram’s camera while recording.

From there, Instagram’s model works through three visible steps, analyzing your clips, identifying the best highlights, and trimming everything down automatically. In one striking example, the tool condensed nearly five minutes of rollerblading footage into an engaging 17-second Reel.

Every decision it makes stays fully reversible too, so you can reorder clips, adjust the cut, or add music before posting anything. Instagram says the First Draft feature eliminates the tedious, frustrating first step of editing, and it can benefit creators with accessibility needs.

Where does this fit into Instagram’s other editing tools?

First Draft lives directly inside Instagram’s existing Reels flow, so you don’t have to jump into a separate app to use it. This feature is quite different from Edits, which is Instagram’s standalone video editor built for more advanced work like frame accurate timelines, storyboards, and AI powered effects.

Interestingly, Instagram hasn’t confirmed whether artificial intelligence actually powers First Draft, so for now it’s best understood as a smart automated tool rather than something explicitly AI branded. In 2026, Instagram has rolled out multiple creator-focused features like Replace Audio and Series, all aimed at making short form content easier to produce. Whether First Draft becomes your new go-to depends on how much creative control you’re willing to hand over, even briefly, to get started faster.

Advertisement

Source link

Continue Reading

Tech

‘Darth Vader’ Wants You to Know He Definitely Supports Flock Surveillance

Published

on

At the San Diego City Council’s most recent budget meeting, Sith Lord Darth Vader had some points to make. Between gasping breaths, Vader praised the council’s recent votes to renew its contract with the surveillance company Flock Safety, which has been making headlines recently as more and more stories of police officers abusing its tools have continued to come out—and as a WIRED review of code for a new AI tool already in use by some police departments has made clear that it is watching far more than just license plates.

“The emperor is a fan of Flock. And we must continue utilizing Flock technologies so that we can follow and surveil the rebel scum as they move from playground to playground, from playground to pool, from pool to gymnasium.”

The recording of the video has gone viral, reaching millions of people on TikTok and Instagram.

The man behind the stunt is Anthony Ralphs, a San Diego-based DJ, event producer, and agriculture consultant who has spent the better part of the past year advocating against the city’s support of Flock’s use by the San Diego Police Department.

Advertisement

Initially, Ralphs thought about keeping his identity hidden behind the Vader mask—a 2004 Hasbro model. But he says that as the video has picked up more and more steam, he decided to “take ownership” of the stunt that put Flock and surveillance technology on the minds and feeds of millions of people around the world.

WIRED spoke to Ralphs about advocacy, crafting a costume, and the role that art can play in civic engagement.

This conversation has been edited and condensed for clarity.

VITTORIA ELLIOTT: How did this start for you? Was this your first city council meeting?

Advertisement

ANTHONY RALPHS: No. My first time attending a city council meeting was sometime last summer or last fall. But really the most consequential city council meeting that I attended, that really got me in the fold of starting to regularly participate in my civic duties, was this issue of Flock. It was my third, or maybe my fourth, city council meeting back on December 9th, and the city council was putting this up to a vote, whether to continue to use Flock.

The story back then was that the city council didn’t exactly know what the technology was. The police were touting it as a force multiplier. There were over 300 people at that vote, the meeting room was completely full, and I would say 95 percent of the people there were saying they didn’t want Flock. But the council voted for it anyway. It was extremely discouraging, and it really lit a fire under me to start going and regularly participating.

I came to the following city council meeting, and I was one of maybe five people in the city council chambers, and there had been close to 300 just a couple of weeks prior. That moment was eye-opening. I realized not everyone can be there all the time to do this, and that’s what we need. We need consistent advocacy. So I took time off of school to be able to do that.

Why did this feel so critical?

Advertisement

Our city’s $120 million deficit. They cut several arts programs, and while they did that, they also provided the $2.2 million for Flock’s contract, even when people in those communities that are being most heavily policed say that they don’t want this. And now, we have the case of Hugo Parra, a local San Diegan who was falsely imprisoned and sat in prison for 30 days due to an erroneous tag by the Flock cameras. [Flock did not immediately respond to a request for comment on this matter.]

Source link

Continue Reading

Tech

Luffu’s New Health and Safety Wearable Isn’t Just for You. It’s for Your Entire Family

Published

on

Wearables like fitness trackers and smart rings traditionally focus on collecting health data at the individual level, but the AI-powered intelligent family care system Luffu is expanding that scope.

Launched by Fitbit co-founders James Park and Eric Friedman in February, Luffu started out as an app where family members could log health information, including medications, diet, doctors’ appointments, vaccine records, symptoms and more. It’s currently in public beta.

Luffu announced on Tuesday that it now has its own screenless wearable, the Luffu Link, and it’s available for preorder in the US for people age 13 and up.

How Luffu links the entire family together

A press release described the Luffu Link as a health and safety band that operates over 4G LTE, meaning you don’t need to have your phone nearby to reach out to your contacts for help or use certain features, like voice logging for medications, meals, mood, symptoms, reminders and notes. It also has built-in GPS, so you can share your location or send a check-in to family members.

Advertisement

When it comes to your personal health, Luffu Link can capture data around sleep, activity, heart rate variability, resting heart rate, calories burned and breathing rate. AI then uses this information to learn patterns and alert users to any changes.

A hand holding a phone showing different family members on the Luffu app.
AI brings all your family’s health data into one app.Luffu

The band features a single button. When you press and hold the button, it allows for voice logging. Press twice to send a check-in to select family members. Any health information you log is organized by AI and then added to the family’s Luffu app, where you can also create a profile for your pet to log their medications, meals and vet visits.

According to the press release, families decide what information gets shared and with whom.

Three phone screens showing different family members and a dog on the Luffu app.
Coordinating care can be accomplished in the Luffu app for both family members and furry friends.Luffu

The Luffu timeline and cost

Luffu is designing its ecosystem around three locations: the phone, wrist and home.

While the app is in public beta, broader US availability is planned for later this year.

As for Luffu Link, those in the US can now preorder it on the Luffu website for $250 and get beta access to the app along with a quick charger and large and small silicone bands. Available in halo gold, sterling blue and onyx black, Link is expected to start shipping in early 2027. After that, the expected price will be $300.

Advertisement
The Luffu Link wearable shown in gold, black and metallic blue.
Designed to be jewelry-like, Luffu Link has a slim metal core with a silicone band. Luffu

Water-resistant up to 50 meters, Luffu Link offers battery life of up to five days, a charging time of around 1 hour, a one-year limited warranty, and compatibility with Android 14 or later and iOS 18 or later.

LTE data service is included with no separate carrier plan needed. However, to use Luffu, a family plan is required. Sold separately and available later this year, the plan covering up to four people costs $20 per month or $200 per year, while the extended family plan for eight people is $30 per month.

Luffu is also developing an intelligent home device — expected to be announced later this year — that allows families to coordinate care across generations and households. In the meantime, the company sent over a preview photo of the device.

The Luffu home device.
What Luffu’s home device will look like.Luffu

The intended audience

A Luffu representative told CNET that some are confusing Luffu Link as a device specifically for aging adults, but it’s intended for the entire family.

“That includes adults [ages] 35 to 60 managing kids, aging parents, pets and their own health; independent, active adults 65+ who want discreet health and safety support without a traditional emergency pendant; and younger or middle-aged adults who may start with health tracking or voice logging and also benefit from safety, location and shared care features,” the representative said.

Co-founders Park and Friedman said they were inspired by their own caregiving experiences to create Luffu and their goal is to give families a single place to easily track their loved ones’ health and safety.

Advertisement

Source link

Advertisement
Continue Reading

Tech

KEF LS LUXE Wireless Speakers Video Review: Is This $4,000 Hi-Fi System Worth It?

Published

on

The KEF LS LUXE are a $4,000 all-in-one wireless speaker system aimed at listeners who want serious hi-fi performance without building a traditional stack of amplifiers, streamers, DACs, and speakers. Although the name may suggest a connection to the smaller LSX II, the LS LUXE are much closer to the LS50 Wireless II in size, capability, and ambition. Their most obvious point of separation is the enclosure, created by renowned industrial designer Ross Lovegrove, who also designed KEF’s radically sculpted Muon loudspeakers. The rounded cabinets give the LS LUXE a far more distinctive visual identity than KEF’s other wireless models and make them as much a design object as an audio system.

Skip the reading? Click play above or watch KEF LS LUXE review on YouTube.

kef-ls-luxe-wireless-speaker-colors-front-angle

Inside, KEF uses its 12th-generation Uni-Q driver array with a 6.5-inch aluminum-cone woofer and 1-inch aluminum-dome tweeter, along with Metamaterial Absorption Technology and the company’s Music Integrity Engine DSP. The LS LUXE also introduce VECO, which monitors driver movement and compensates in real time, along with KEF’s Current Drive amplification technology designed to further reduce distortion. Amplification totals 280 watts of Class D power for the woofers and 100 watts of Class A/B power for the tweeters, while connectivity includes HDMI eARC, optical, coaxial, 3.5 mm analog, Ethernet, subwoofer output, AirPlay, Google Cast, Spotify Connect, TIDAL Connect, Bluetooth 5.3, and internet radio. There is no phono stage or RCA analog input, so turntable users will need an external phono preamp and, depending on their setup, an adapter.

kef-ls-luxe-primary-speaker-connections-closeup

Do the LS LUXE improve on the LS50 Wireless II? In some respects, yes. KEF has added newer driver-control and amplification technology, the Lovegrove-designed cabinet is far more visually ambitious, and the sound is exceptionally clean, neutral, detailed, and controlled, with impressive bass speed and very low perceived distortion. But the jump is not transformative enough to make the LS50 Wireless II obsolete, especially when you consider the roughly $1,000 premium and the inevitable diminishing returns at this level. Buyers are paying for both the engineering and the industrial design. For a closer look at how the LS LUXE perform with music and movies, and whether they can really justify that $4,000 asking price, watch our full video review on YouTube.

Advertisement
kef-ls-luxe-wireless-speakers-front-back-mineral-white

The Bottom Line

The KEF LS LUXE earn our Editors’ Choice designation because they combine exceptional ease of use, superbly clean and controlled sound, and a genuinely distinctive industrial design in a compact wireless system that looks unlike anything else in the category. Ross Lovegrove’s sculpted enclosure, KEF’s latest driver-control technologies, HDMI eARC connectivity, and excellent overall refinement make the LS LUXE one of the most complete premium wireless speaker systems we’ve tested. They are also unusually versatile, working equally well as a serious music system or an elegant alternative to a conventional TV audio setup.

The catch is value. Buyers are clearly paying a premium for Lovegrove’s design, the compact form factor, and KEF’s newest engineering, but the sonic improvement over the LS50 Wireless II is not large enough to make that model suddenly irrelevant. Klipsch’s The Sevens II remain the stronger value, particularly for TV use, at half the price. But if design matters as much as sound and you want a compact, beautifully finished, genuinely high-end wireless system that requires very little fuss, the LS LUXE occupy a space that few competitors can match.

Pros

  • Incredibly clean, precise sound
  • Easy to setup and use
  • Gorgeous industrial design
  • HDMI eARC for connecting to TVs
  • Digital connectivity options
  • Included high quality remote control
  • Supports Hi-res audio up to 192 kHz/24-bit
  • Optional finishes

Cons

  • Doesn’t glow up lesser recordings
  • No RCA or Phono Inputs
  • No volume/play controls on speaker itself

Our Ratings

★★★★★★★★★★ Sound Quality

★★★★★★★★★★ Build Quality

★★★★★★★★★★ Usability

Advertisement

★★★★★★★★★★ Value

Where to Buy

Advertisement. Scroll to continue reading.

Source link

Advertisement
Continue Reading

Tech

Prompt injection ranks No. 1 with OWASP and No. 12 in the incident record. The attack itself is invisible to a scan.

Published

on

A CISO who sees a low CVE count and deprioritizes prompt injection is reading the scoreboard wrong. Prompt injection has held the No. 1 spot on the OWASP Top 10 for LLM Applications for three consecutive years. When two leaders of that list checked it against 6,639 labeled real-world incidents, it came back at No. 12. The drop measures visibility rather than danger, because the attack operates where a vulnerability scanner cannot see it.

That finding belongs to Kyriakos “Rock” Lambros and Steve Wilson, two leaders of the OWASP Top 10 for LLM Applications project, who published it on arXiv on August 18 with the disclaimer attached. The analysis is exploratory, not peer reviewed, and not the official OWASP release, and the authors state it does not supersede the official list or its process.

The machinery behind it is real: 7,714 LLM security incidents from CVE, GitHub Security Advisories, OSV, and the AIAAIC AI-harm database, 6,639 of them labeled against a 20-entry taxonomy, and a Bayesian model that corrects each count for classifier error before setting the data-driven ranking beside the expert vote.

The comparison found no statistically detectable agreement between expert judgment and the public incident record. Cohen’s kappa comes in at 0.20 with a 90% interval running from negative 0.16 to 0.57. “The interval crosses zero, so we cannot rule out that the two rankings agree only by chance,” they write. “The honest bottom line: weak agreement, not confirmation.”

Advertisement

Lambros, co-lead of the OWASP GenAI Security Project Top 10 for LLM Applications and director of AI standards and governance at Zenity, put the finding in evidentiary terms in written answers to VentureBeat. “We had two ways of measuring the same risk, expert judgment and the public incident record, and they disagree with each other. Neither one is the truth,” Lambros said. “Two witnesses are contradicting each other, and we can’t tell you which one is lying.”

The attack chain a scanner never logs

The gap is structural. Prompt injection hides instructions inside the content a model reads, anything from a log entry to a support ticket to a document pulled back by retrieval. The agent then makes the tool call the attacker wanted, using credentials it legitimately holds. Nothing in that chain is a product defect, so the attack leaves no CVE behind for a scanner to find.

The defenses that catch it are adversarial tests against the deployed system and hard caps on what the agent can reach, so a fooled model cannot touch anything expensive. The same logic argues for funding agent memory and MCP tool boundaries now, on architecture, rather than waiting for advisory volume that will always arrive a cycle late.

The first control Wilson would deploy

Wilson, Chief AI and Product Officer at Exabeam and project co-lead for the OWASP Top 10 for LLM Applications, named the control he would deploy first against exactly that chain, an agent that reads an attacker’s payload in a log file, treats it as an instruction, and rewrites DNS with a valid credential, in written responses to VentureBeat.

Advertisement

“The first thing I’d do is put an authorization gate outside the model: the agent can propose the exact DNS change, but it cannot grant itself the authority to make it,” Wilson said. “Security rules written inside prompts may shape the model’s behavior, but they are still suggestions to the model, not enforceable security controls.”

The gate has a price, and Wilson states it plainly. “The tradeoff is that the agent loses the ability to improvise arbitrary, high-impact infrastructure changes on its own, while retaining autonomous investigation and routine, bounded remediation,” he said.

Why the No. 1 risk looks small in the record

“Prompt injection is the best-understood LLM attack, and deployed systems defend against it actively,” the authors write, and they compress the whole divergence into one sentence. “Experts rank it first because the attack surface stays enormous even when the defenses mostly hold; the data sees the successes that got through.”

Wilson has watched the gap from both sides of it. “Incident data is incredibly valuable, but it is inherently backward-looking and notoriously tricky to interpret,” he said. “It tells us what was observed, recognized, classified, and reported. It does not necessarily tell us what is most dangerous in the systems people are building right now.”

Advertisement

He compares prompt injection to “death and taxes” and, increasingly, to “a law of physics for LLM systems,” because one model is being asked to interpret trusted instructions and untrusted content at the same time.

Better defenses have not closed the case. “A control that works 99% of the time is not sufficient when the failure case gives an attacker meaningful access. And, frankly, I don’t think we are at 99%,” Wilson said. “The durable answer is not believing we can perfectly screen prompt injection out of existence. It is designing systems with the assumption that prompt injection will occur, understanding why it works, and limiting what an attacker can accomplish when it does.”

A low advisory count can mean the defenses are working. It can just as easily mean nobody has looked, and the public record cannot tell a security team which one it is.

The attempt volume is documented. CrowdStrike’s 2026 Global Threat Report found adversaries injected malicious prompts into legitimate GenAI tools at more than 90 organizations in 2025, stealing credentials and cryptocurrency, under a section titled “Prompts are the New Malware.” The telemetry shows pressure on the attack surface without proving defenses produced the No. 12 placement, but it is the pattern the mechanism predicts.

Advertisement

The gap runs the other way too, and further

Prompt injection is the headline case, and misinformation is the bigger one.

The expert vote puts misinformation at No. 13, while the incident record places it at No. 2. The paper calls it “the widest disagreement between the two witnesses” and reports that its concordance flag “puts the probability that the two signals disagree at 99 percent.”

The authors do not treat their own data as the winner. On misinformation they note the corpus “carries a large volume of deepfake and AI-generated disinformation,” records that often “describe harm produced by an AI rather than a vulnerability inside an LLM.” The authors call it the entry the record most disputes, stopping short of concluding the experts got it wrong.

Where “too new to measure” runs into the CVE record

The two brand-new taxonomy entries sit at the sharpest end. Persistent memory poisoning lands at expert No. 4 and incident No. 16, MCP tool interface exploitation at expert No. 7 and incident No. 16, each with an incident interval of 6 to 20 that spans most of the taxonomy.

Advertisement

Public 2026 CVEs exist for both. On MCP tool interfaces, the Azure Data Explorer MCP Server carried KQL injection, and the CVE record describes it allowing “an attacker (or a prompt-injected AI agent) to execute arbitrary KQL queries against the Azure Data Explorer cluster,” scored 8.3 High. Kong’s Konnect MCP Server shipped an indirect prompt injection that lets a remote attacker steer the server into executing unintended API requests, the exact failure the MCP entry names.

Agent memory has its own record. An agent harness, Ruflo, exposed unauthenticated MCP bridge endpoints that let a network attacker obtain a shell, read provider API keys, and poison the learning store, rated 10.0 Critical.

The record is so thin and uncertain that the model cannot place either entry within 14 rank positions. A team waiting for advisory volume to justify a control on agent memory or an MCP tool boundary would still be waiting while the CVEs accumulate at Critical and High.

Lambros makes the budget case in operational terms. Poisoned memory “doesn’t announce itself,” he said. It looks like a procurement agent told once that invoices from a given supplier under $50,000 clear without a second signature, and because the agent remembers, every approval after that looks like the process working. “Nobody files an advisory for that, because nobody knows it happened. A count of zero is measuring your blindness, not your safety.” The argument he says a CFO will sign off on is timing, since memory and tool permissions get wired into these systems once, early, and everything else sits on top of them. “Build it in now and it’s a rounding error. Come back in two years and you’re re-architecting and re-training your systems.”

Advertisement

The authors flag their own measurement problems first

The expert side is thin. “The expert signal is a practitioner survey: about 29 respondents scored each candidate risk on importance,” the authors write. Twenty-nine votes set the ranking that carries three-quarters of the published list’s weight, the compression point for OWASP’s more than 25,000 community members.

On the data side, the classifier is the weak joint. Precision “varies sharply across entries, from 93% (LLM01, LLM03) down to 13% (LLM08),” four entries fall below 50%, and the base classifier “never predicts ‘out of scope’ and files every incident into some category, including the roughly 38% of the gold set that belongs in none.”

The authors name the central limitation themselves. One reviewer adjudicated all 1,200 gold-set incidents and overrode the model consensus on 553 of them. “A single annotator cannot measure inter-rater reliability,” they write. “The single-author gold set remains the central limitation.”

Lambros lays the weak kappa at the feet of the taxonomy itself. “That number is telling you about our categories, not about our experts,” he said. When the people who wrote a taxonomy cannot reliably sort incidents into it, he argues, “a weak score on the ordering of those buckets is a fact about the buckets.”

Advertisement

A better classifier will not fix the disagreement. A pre-registered bake-off of four frontier models produced no winner. None beat the incidence floor’s balanced accuracy of 0.863, and a ground-truth check left the floor’s ordering in place at a Spearman correlation of 0.918. The authors published the engine and artifacts on GitHub for anyone to rerun.

The robustness result tested only one side of the gap. Every check behind the abstract’s word “robust” runs on the incident side, showing the incident-derived ranking stays put when the labeling machinery changes, and none of it touches the 29-vote survey. A board that hears “robust” will assume validated, yet the record supports only stable.

What the published list did with this

OWASP shipped the GenAI LLM Top 10 2026 on August 4, the first edition to fold incident data into the ranking, weighting the practitioner vote at 75% and the incident corpus at 25%. Prompt injection stayed at No. 1, misinformation moved up two places, excessive agency climbed from No. 6 to No. 3 as the entry where the two signals agree most clearly, unbounded consumption rose four spots to No. 6, and improper output handling fell from No. 5 to No. 10, the largest drop.

Wilson declines to defend the blend as arithmetic. “There is nothing magical about a 75/25 weighting,” he said, “or about reversing it to 25/75. The value of the data wasn’t that it gave us a mathematical answer; it changed the conversation.” The excessive agency entry is where that conversation landed hardest for him. “If I were a CISO evaluating a new agentic deployment today, Excessive Agency is where I would start,” Wilson said.

Advertisement

Lambros would go further next cycle, a view he flags as his own and separate from the working group. The blend hands the same 25% incident weight to every category, while the hand-checked classifier precision runs from roughly nine in 10 on prompt injection and supply chain down to roughly one in eight on vector and embedding weaknesses. A quarter of the weight on the first rides on something solid, he argues, and the same quarter on the second rides on noise. “The ratio should track how well we actually measure each category,” Lambros said.

Why this lands now

Ivanti’s 2026 State of Cybersecurity research found 87% of security teams call adopting agentic AI a priority and 77% report at least some comfort letting AI act without human review. Teams are signing off on agent autonomy while the expert ranking of what can go wrong with those agents shows no statistically detectable agreement with the incident record.

What to do with this on Monday

The behavioral change is narrow and it is the whole point.

  • Use the OWASP LLM Top 10 as a coverage map, not a queue. The rank positions carry 29 votes and a corpus whose own authors call the agreement weak, so build your own priority order from your own exposure: production reach, breach-notification data, and controls that have actually been tested. Lambros draws the funding line the same way. “I’d prioritize spend where the expert vote and the incident record point the same direction, because that’s two independent witnesses agreeing,” he said. “Where they split, stop letting the ranking allocate your money and go look at what your own systems are doing.”

  • Log what your AI systems are actually doing, field by field. The prompt that went in, what came back out, the documents pulled to build the answer, the tools called and the arguments passed to them, and the model’s confidence score on every response. Confidence is the field Lambros would fight for, because most security leaders do not realize it is measurable, and it is where the attack surfaces. “A model running on a poisoned instruction doesn’t act broken. It acts certain,” he said. “Certainty is what your monitoring treats as a healthy system.” The cost is a sprint or two of engineering. The constraint is a person, because a SIEM does events and these are trends. “Somebody has to analyze those trends every week and say whether a drift means anything, and most security teams have nobody who can.”

  • Stop expecting scanner output to reproduce the Top 10’s order. Scanner findings live on the incident side of the gap, counting what got disclosed rather than what a deployed system should fear, and the classifier bake-off shows a smarter model does not close that distance. The test that sees prompt injection is an adversarial one run against the live system, paired with Wilson’s authorization gate so the change an injected agent proposes is never the change it can execute.

  • Fund the thin-record categories on architecture, not incident volume. Agent memory and MCP tool boundaries sit at expert No. 4 and No. 7 with incident intervals spanning most of the taxonomy, and the CVEs that do exist are landing at High and Critical. Kayne McGladrey, an IEEE senior member who advises enterprises on risk, put the funding logic bluntly in an interview with VentureBeat. “Anything that seems to have a cybersecurity flavor is generally put into the cybersecurity risk category, which is a complete fiction,” McGladrey said. “They should be focused on business risks, because if it doesn’t affect the business, like a financial loss, then nobody’s going to pay attention to it, and they will not budget it appropriately.” A rank number from a 29-person vote is a weaker budget argument than the business system the agent touches.

  • Steal McGladrey’s baseline test for the AI systems themselves. “If you wouldn’t expose your database to the public internet without identity and access controls, why would you do that for your AI model?” he said in CSO Online’s analysis of 2026 breach costs.

The board question for the next meeting is short. If our AI risk ranking came from a 29-person vote and a corpus that disagrees with it, what are we actually using to decide which controls get funded next year?

Advertisement

Source link

Continue Reading

Trending

Copyright © 2025