Connect with us

Tech

Monitor Audio’s Silver Series 8G reinvents its best-selling speaker line-up

Published

on

Monitor Audio has unveiled the Silver Series 8G, a complete redesign of its best-selling stereo speaker range with seven new models aimed at both stereo listeners and home cinema enthusiasts.

The new range introduces upgraded driver technology, a redesigned tweeter, a new port design and improved cabinet bracing. Monitor Audio positions the Silver Series 8G as its most advanced Silver range yet with speakers designed to scale from compact 5.1 systems through to full Dolby Atmos 7.1.4 setups.

At the heart of the range is the latest version of Monitor Audio’s C-CAM driver technology, now paired with RST III cone geometry. The company says the updated design increases rigidity while reducing unwanted resonances, the aim is to deliver greater accuracy and control across the frequency range.

The C-CAM Gold Dome tweeter has also been redesigned. It now features a vented rear chamber, a new shorting ring and the company’s UD Waveguide II. These are intended to improve clarity, dispersion and consistency at higher frequencies.

Advertisement
Monitor Audio Silver 8G AV speakerMonitor Audio Silver 8G AV speaker
Image Credit (Monitor Audio)

Advertisement

Another significant change is the new Silent Flow port technology. Inspired by the structure of an owl’s feathers, the port uses an enlarged diameter and revised flare geometry to smooth airflow. Monitor Audio claims this reduces turbulence while helping the speakers deliver deeper, cleaner bass.

The cabinets have also received structural upgrades, including Through-Bolt Driver Bracing, which secures the drive units directly to the rear of the cabinet. New outrigger feet, inspired by the flagship Platinum Series 3G, provide a wider base and adjustable levelling. Furthermore, integrated rubber feet can be used instead of spikes on hard floors.

The seven-model range comprises the Silver 500 8G, Silver 300 8G, Silver 100 8G, Silver 50 8G, Silver On-Wall 8G, Silver Centre 8G and Silver AMS 8G. They are available in High Gloss Black, Satin White and European Walnut finishes.

Pricing for the Silver Series 8G is as follows:

Advertisement
Silver 500 8G  £2650 €3100 $3300
Silver 300 8G £2000 €2400 $2600
Silver 100 8G  £1500 €1800 $2000
Silver 50 8G £800 €1000 $1050
Silver On-Wall 8G £450 €550 $575
Silver Centre 8G £800 €1000 $1050
Silver AMS 8G  £900 €1100 $1200

Advertisement

Source link

Continue Reading
Click to comment

You must be logged in to post a comment Login

Leave a Reply

Tech

From furniture to warships: this US factory 3D-prints drone boats using robot arms

Published

on


  • A Florida company has moved from printing furniture to manufacturing military drone boats
  • Haddy built its latest unmanned boat using giant robotic 3D printers
  • The company completed an earlier autonomous vessel project in nine days

Haddy, a digital manufacturing company in St. Petersburg, Florida, has produced a military drone boat using large-format robotic 3D printing.

The company announced the vessel, called the TF-179 Drone Boat, but did not name the customer behind the project.

Source link

Advertisement
Continue Reading

Tech

This R-Rated Film Studio Wants to Be the HBO of AI

Published

on

Demand for artificially generated smut is surging. But the quality of the visual content being created and shared—like on companion apps, where adult performers are selling their likeness to give fans a 24/7 experience, WIRED reported in March—isn’t exactly movie-level caliber.

Rogue Studio, a cinematic AI-generator platform built around adult visual storytelling, is trying to change that. Launching today, the studio, which bills itself as a “playground for creative ethical mischief,” plans to deliver industry-level video production without creative and sexual restrictions. Its marquee product is Rogue 1.0, an AI video-generation tool for making uncensored adult content that is comparable to other pro-level models Hollywood is using.

“There really isn’t a place for a sophisticated creator to go make content that’s rated R. Either they go to these skeezy NSFW sites that are built on outdated open-source tech, or they try to do high-quality stuff but they’re censored up the gills,” a Rogue Studio cofounder, who goes by Mr. Rogue, tells WIRED. “It’s a very fine line between being a sophisticated platform for adults versus becoming just a porn site. And we do not want the latter. We want to do what HBO did for television.”

Rogue, which has subscription tiers ranging from $29 to $299, offers all the safe-for-work frontier models within its system—similar to Higgsfield, it includes NanoBanana, Seedance, Kling 2.0, and others—in addition to its own proprietary models for generating adult-oriented images and video. Using the platform, you can employ multiple elements from several different AI-generator models when drumming up ideas.

Advertisement

Let’s say you have a concept for a freaky sci-fi horror short. You could potentially start ideating on a zombie character using Midjourney, then use Rogue’s AI tools to add more high-definition uncensored features to the character.

According to its terms of service, Rogue prohibits user-generated images or videos for training, and it strictly forbids the creation or distribution of nonconsensual deepfakes, child sexual abuse material, unauthorized depictions of public figures, and other illegal or abusive imagery.

In an effort to prevent bad actors from using the platform to make these prohibited visuals, Rogue has three layers of human moderation. The first is at the prompt level, which searches for keywords and flags them for possible account suspension if it is determined that there is clear illegal intent. The next layer is inside the system, called the intention layer. Bad actors might try to outsmart moderators by slightly misspelling a real actor’s name to generate a deepfake image, so moderators will reread the prompt to assess the user’s intent. The final layer analyzes the image or video created before it is shared with the user. For example, if a user prompts “famous blonde singer” and the image generated looks like Taylor Swift, the image won’t get delivered.

Mr. Rogue tells WIRED that he and his fellow cofounder, both of whom are movie producers, have decided not to disclose their names because of the “real risk” involved in being associated with creating erotic-themed work. “Because of the climate right now in Hollywood and working in AI, you have got to be very careful,” Mr. Rogue says. (WIRED verified that the pair of founders do, in fact, work in Hollywood; they have produced movies with a combined box office of over $2 billion.)

Advertisement

In the wake of the 2023 Hollywood strikes, where labor unions fought to secure guardrails against the unapproved digital likenesses of actors appearing in films and TV, leading studios have been assessing ways to incorporate AI into productions with the goal of lowering expenses.

The friction has led to AI being seen as unfavorable and has created a stigma around it. Rogue’s cofounders know this but don’t want to shy away from the tech’s creative potential.

Source link

Advertisement
Continue Reading

Tech

Jorts Freezer Pocket Keeps Ice Cream Sandwiches Cold on Hot Days

Published

on

Jorts Freezer Pocket Cooler Ice Cream
Summer afternoons turn ice cream sandwiches into sticky messes in minutes. One maker decided that ordinary pockets would no longer do. Brendan Carberry built a battery-powered storage case that clamps straight onto a pair of jorts and chills the treat inside while the heat pumps out the other side.



He calls it the Jorts Ice Cream Sandwich Pocket. Three heat pumps are situated in the center. The Peltier elements move heat from the cold side (which is attached to the sandwich) to the hot side (which is protected by some large, chunky heatsinks). A little battery that fits in the pocket provides the power needed to run everything. And there’s a little remote that allows you to adjust the temperature till it’s exactly right. The whole thing stays chilly on the inside (the sandwich keeps firm) and warm on the outside (so the rider’s legs don’t freeze), which is ideal.

Sale


Instant Pot InstantChill Ice Cream Maker with Built‑In Compressor, No Pre‑Freezing, Real Ice Cream in…
  • NO PRE‑FREEZING, ICE CREAM IN MINUTES*: Built-in compressor and cold plate system rapidly freezes and churns ultra-smooth fresh ice cream, gelato…
  • BUILT‑IN COLD PLATE FOR FASTER RESULTS: Churn in the bowl or pour directly onto the cold plate for even faster freezing. Roll, or scoop—your…
  • 6 ONE-TOUCH PROGRAM MODES: (6) precision pre-set programs deliver the ideal balance of speed and timing for perfect results every time. Make Ice…

Jorts Freezer Pocket Cooler Ice Cream
Building it began with a 3D print of the sliding case, which was a basic ice cream sandwich case, nothing spectacular. Some magnets and matching sewing tabs were inserted in there to clip onto the fabric of the jorts. Carberry put the tabs onto his jorts, removed them again, and sewed the magnets in place. He then installed the three Peltier units in the casing, applied thermal goo to the heatsinks, and ran the wiring. Once the battery and control board were attached, it was time to turn on and go.

Jorts Freezer Pocket Cooler Ice Cream
The completed device resembles a swollen cargo pocket with a cold container inside. When you click the magnet shut, it locks in place. On the hottest days, the sandwich remains hard enough to consume without dribbling down your shirt or melting all over. He designed this precisely for when it’s a million degrees outdoors and all you want is something cool to eat, without having to haul about a large cooler or worry about your ice packs getting mushy after a few minutes of walking.

Jorts Freezer Pocket Cooler Ice Cream
He had previously experimented with building small cool boxes using the same heat-pump concept, which taught him how to cope with waste heat. He’d also built some other coolers that showed him the value of good airflow and large heatsinks, so this pocket-sized version still provides some useful cooling, just long enough for a brief outing or a trip to the park. Battery life is limited, and these solid-state coolers lack the power of a compressor, but they still work.

Advertisement

Source link

Continue Reading

Tech

OLEDs Have Gained Brightness, Not Burn-In Resistance

Published

on

OLED displays solve many of the problems suffered by LC displays, including color fidelity, dynamic range and power usage. That said, especially in the early days OLED gained a reputation for dim screens, short lifespans and burn-in. Over time better organic dyes were developed, along with burn-in prevention methods that have made OLEDs much closer to LCDs in terms of longevity. In a recent comparison between OLED TVs by RTings it’s however clear that between 2017 and 2023 there haven’t been any major advances beyond bumps in brightness.

The relatively dim screens were a major problem, as they prevented OLEDs from displaying HDR content. This issue has been well and truly addressed, as confirmed by RTings’ testing, but after an over 10,000 hours stress test that simulates about 10 years of regular use at maximum SDR brightness, especially static elements like the CNN TV banner happily burned in even on the newest models with all burn-in prevention measures enabled.

Here the biggest take-away is probably that even if the expected panel lifespan at full brightness is still the same, this higher brightness budget means that you can gain some lifespan by cranking the brightness way down. It’s also essential to keep features like pixel refresh cycles enabled, as demonstrated by [Hardware Unboxed] and their abuse of a QD-OLED monitor where a worst-case 6,000 hour stress-test managed to create some impressive levels of burn-in from uneven subpixel wear.

Advertisement

Source link

Advertisement
Continue Reading

Tech

S’pore has become a two-speed economy. AI boosts GDP, but some industries are left behind.

Published

on

Disclaimer: Unless otherwise stated, any opinions expressed below belong solely to the author. Data sourced from Singapore’s Ministry of Manpower.

Singapore’s economy is having an exceptionally strong year. GDP grew 5.9% year-on-year in the second quarter of 2026, after expanding 6.3% in the first. This led the Ministry of Trade and Industry (MTI) to raise its full-year forecast to 4.5-5.5% – up from the original 1.0% to 3.0%.

The main reason is Artificial Intelligence (AI).

Global spending on AI infrastructure is boosting demand for semiconductors, manufacturing equipment, cloud services and other technology-related activity. Singapore happens to be extremely well positioned to benefit.

Advertisement

But the gains are becoming increasingly concentrated.

Source: Economic Survey of Singapore Q2 2026./ Ministry of Trade & Industry

In Q2, manufacturing grew 12.5%, with electronics output surging 33.8% and precision engineering rising 19.3%. Wholesale trade expanded 8.3%, while finance and insurance grew 6.2%.

Together, manufacturing, wholesale trade and finance accounted for around three-quarters of Singapore’s GDP growth during the quarter. Elsewhere, things looked rather different.

Retail grew by just 1%, accommodation 2.2% and professional services 2.4%, while embattled F&B contracted by 1.5%.

Singapore increasingly looks like a two-speed economy.

Advertisement

Productivity gap

The difference is even clearer in productivity. Value added per hour worked increased 15.4% in wholesale trade, 9.4% in information and communications, 7.6% in manufacturing and 5.1% in finance.

Across outward-oriented industries, productivity rose 6.9%. Among domestically oriented industries, it fell 0.1%.

This explains why GDP can grow close to 6% without everybody feeling that the economy is booming. 

A semiconductor factory can increase output dramatically without hiring thousands of additional workers. The same is true of cloud computing, finance or wholesale trade. Restaurants and shops cannot scale in quite the same way.

Advertisement

Salaries

Over time, productivity growth is what allows wages to rise sustainably. That should benefit Singapore. But if it’s concentrated in one corner, it may widen differences between workers, creating a headache for the government, as some of the people benefit greatly while others see their incomes slide behind in real terms.

MTI research on AI adoption found that initial employment gains at companies using AI were concentrated among higher-earning local workers, mid-career employees and skilled foreign professionals.

Only as firms developed deeper AI capabilities did the benefits begin spreading more widely.

This suggests the early winners are likely to be engineers, semiconductor specialists, software developers, data professionals and workers in related business services.

Advertisement

But someone working in retail or F&B is participating in a very different economy.

There is already evidence that productivity is rising much faster than labour costs in some of the booming sectors. Unit labour costs fell 7.9% in manufacturing and 3.7% in wholesale trade in Q2. 

That creates room for higher wages—but there is no guarantee the gains will be distributed evenly.

What happens if the AI boom ends?

This is a risk that was flagged by MAS, even as GDP figures should have the country celebrating. If AI-related investment is now responsible for a large share of Singapore’s growth, what happens if the boom ends?

Advertisement

Singapore would be exposed across several sectors at once, and the threat is not only a freeze at the current level of demand, but a dramatic contraction which could result in mass job losses.

Lower AI spending would weaken semiconductor demand. That would hit electronics manufacturing and precision engineering. Lower trade volumes would affect wholesale and logistics, while technology and financial services could suffer from weaker investment and asset prices.

Could that push Singapore into recession? Yes—if the reversal were severe enough.

That does not mean an AI downturn would automatically cause one. Construction remains strong, domestic consumption continues to grow, and Singapore’s economy is diversified.

Advertisement

But current growth is unusually concentrated.

When three sectors account for roughly three-quarters of quarterly GDP growth, losing momentum in those industries can change the headline numbers very quickly.

Singapore is benefiting enormously from the global AI investment boom.

But the scale of its growth also shows how exposed it is to downside risks if the bubble suddenly pops.

Advertisement
  • Read other articles we’ve written on Singapore’s current affairs here.

Featured Image Credit: Leo Heng/ Unsplash

Source link

Advertisement
Continue Reading

Tech

The Pixel 11 gives me three reasons to ditch my iPhone 17, and I’m afraid they’ll work

Published

on

As someone who reviews smartphones for a living, upgrading every year isn’t optional. It’s part of the job, and as an iPhone user, it’s an easy call. This year, I expected to move from my baseline iPhone 17 to Apple’s next baseline model, the iPhone 18. Except Apple might not even launch it this September.

Pegatron, one of Apple’s suppliers, indicated on an August 12 earnings call that the standard iPhone 18 isn’t coming next month (via Deccan Chronicle). Only the Pro, Pro Max, and the rumored foldable “Ultra” are expected, with the base model pushed to early 2027. That alone would derail my plans. But it’s just one domino in a longer line for Apple: the Pixel 11.

Pixel 11 vs. iPhone 17: It doesn’t feel like a tie anymore

iPhone 17 Pixel 11
Price $799 $899
Display 6.3″ OLED, 120Hz ProMotion 6.3″ OLED
Chip A19 Tensor G6 (2nm)
RAM 8GB 12GB
Base storage 256GB 256GB
Battery 3,692 mAh 4,985 mAh
Rear cameras Dual 48MP (main + ultrawide) Triple: 48MP main, 13MP ultrawide, 10.8MP 5x telephoto

Here’s the kicker: the Pixel 11 isn’t even the cheaper alternative anymore. Google just increased the starting price to $899 (up from $799), making this a tougher sell, not an easier one.

I’m low-key rooting for the Pixel 11 this year, and the iPhone 18 delay just gave me one less reason to talk myself out of it. The other three have more to do with what Google is offering on its entry-level flagship, starting with the cameras.

This is where the Pixel 11 starts pulling me away

Apple’s iPhone 17 offers a dual 48MP camera array on the back. Although it captures excellent pictures, even if a candle is the primary light source in the frame, the camera setup lacks a dedicated telephoto sensor. The Pixel 9 didn’t have one either, but Google broke that pattern with the Pixel 10, and the Pixel 11 keeps it going with a proper triple system.

I’m talking about a 48MP main, 13MP ultrawide, and a 10.8MP telephoto with 5x optical zoom. Yes, I’d take a hit on the selfie camera, as there’s no match for Apple’s Center Stage experience yet, but since I capture one selfie for every 100 or 150 photos from the rear camera, the math is pretty clear.

Advertisement

The time I spent with the base Galaxy S26 earlier this year simply rewired my expectations from my daily driver. A dedicated telephoto camera puts more emphasis on the subject, bringing out the texture of skin, the fine details on a flower petal, the arch at the top of a building, or the person standing on the other side of the street. It simply feels more intentional.

iPhone 17 Pixel 11
Main 48MP, f/1.6, sensor-shift OIS 48MP, f/1.7, OIS
Ultrawide 48MP, f/2.2, 120° FOV 13MP, f/2.2, 120° FOV
Telephoto None — 2x “optical-quality” zoom is a digital crop from the main sensor 10.8MP, true 5x optical zoom, OIS
Selfie 18MP Center Stage 10.5MP, 95° FOV
Max zoom 10x digital 5x optical, up to 30x Super Res Zoom
Video 4K Dolby Vision 4K, including new 4K portrait video

Google’s software pitch is harder to ignore

Then there’s the software side, which caught me off guard. Android 17 didn’t chase a flashy redesign this cycle. Instead, Google spent its energy sanding down the friction, and the rebuilt Android Switch tool is the clearest proof. It’s wireless now, built directly into both iOS and Android with no separate app required.

The best part is that it now migrates passwords, passkeys, Wi-Fi credentials, alarms, even full text threads. That matters to someone like me who’s hesitant to switch. Task automation via Gemini Intelligence is still an open question for me, but it’s intriguing enough that I want to learn its real-world use cases.

Even otherwise, there are too many Google AI features to explore, especially the ones related to photography. The one that has me sold is something called Magic Capture, which analyzes around 400 shots to determine the best picture and video from the scene you point to. Then there’s Creator Suite, which basically includes built-in AI video editing tools.

Apple Intelligence (iOS 27) Gemini Intelligence (Android 17)
Assistant Siri AI, still rolling out, routes complex queries through ChatGPT and, per some reports, Gemini itself Gemini, native to the phone from day one
Hardware requirement A17 Pro or newer Flagship Tensor chip, 12GB RAM minimum
Writing Writing Tools: rewrite, proofread, summarize Gboard Rambler: cleans up natural speech into text
Camera AI Genmoji, Image Playground Magic Capture, Creator Suite
Screen understanding Visual Intelligence for screenshots and camera Android Halo, persistent AI status in the status bar
Translation Live Translation in Messages and FaceTime Live Translate, real-time voice and video

The battery could be the Pixel 11’s sleeper advantage

Finally, the physically larger battery could solve one of my biggest complaints with the iPhone 17. Google has crammed a 4,985 mAh battery into the base Pixel 11. Combined with the new Tensor G6 (2nm) chip, which the company claims is 20% more power efficient than the previous chip, the Pixel 11 might outlast the iPhone 17

Google’s own claim has gone up from “24+ hours” battery life on the Pixel 10 to “30+ hours” on the Pixel 11. I’m not buying those numbers at face value, but I’m eager to take the Pixel 11 on a spin and check the screen-on time I get between chargers.

Advertisement

This isn’t me pretending the iPhone 17 suddenly became a bad phone for me. In fact, its video remains excellent, the display is still fantastic, performance is more than enough for everything I do, and the Continuity features I get with the MacBook Air M5 make the two devices feel almost inseparable. 

I’d wait for more hands-on experience before making my decision

However, for the first time, those advantages aren’t as appealing as what I could get with a Pixel upgrade. For me, that justifies the $100 premium over the iPhone 17. You’d argue that I can stop crying about it and go with the iPhone 18 Pro instead, but I simply don’t want to spend around $1,200 on my smartphone.

I’m not abandoning iPhone. I’ve just never had this many real, specific reasons pulling me toward the other side at once. And this year, the Pixel 11 is making the better argument. I’m waiting for more hands-on experience with the device, which will actually help me make up my mind before Google gets money from my wallet. 

Source link

Advertisement
Continue Reading

Tech

An AI broke Snowflake’s code. Then another AI agent exploited it

Published

on

Security

Don’t worry, this one was via a bug bounty program

An AI broke Snowflake’s code; then another AI, an attack agent, autonomously found the bug, exploited it, and extracted credentials without human intervention.

Luckily, this wasn’t yet another case of rogue AI agents doing evil things. It was a sanctioned bug hunt, conducted through Snowflake’s HackerOne vulnerability disclosure program, and Snowflake fixed the flaw the same day Wiz reported it and rotated the affected credentials the following day.

Advertisement

Wiz’s red agent, an AI-powered autonomous attacker designed for offensive security, found the GitHub Actions workflow flaw during a routine scan of public repositories on June 23. The script injection vulnerability existed in snowflakedb/snowflake-connector-net, and it allowed an unauthenticated user to execute arbitrary commands within a GitHub Actions runner by opening a GitHub issue with a specially crafted title.

And it turned out an AI had inadvertently injected the bug into the code five days earlier.

GitHub Copilot Autofix, an AI coding assistant, co-authored the commit on June 18, and it introduced a script injection bug in run: blocks by removing the repository’s existing sanitized input pattern and replacing it with direct string expansion in a shell script.

“We crafted an issue title that, after template expansion, breaks out of the echo string and exfiltrates the Jira credentials via an out-of-band callback,” Wiz’s head of threat exposure Gal Nagli said in a Monday blog. 

Advertisement

These credentials gave Wiz read access to Snowflake’s engineering, security compliance, and bug bounty tracking projects.

Wiz reported the workflow vulnerability to the cloud data platform on June 23, and Snowflake patched it the same day. It also revoked and rotated the Jira token, and confirmed, via audit logs, that Wiz was the only third-party to access the endpoint during the five-day exposure window. 

The disclosure “was immediately investigated and remediated, and our investigation found no evidence of unauthorized access,” a Snowflake spokesperson told The Register. “We are working together with Wiz to share these learnings with the broader industry to encourage widespread adoption of these security best practices.”

Wiz, for its part, deleted all of the data it accessed during the vulnerability research and proof-of-concept exploit testing, and told us that this incident proves human code review isn’t sufficient to quickly detect vulnerabilities –  especially as developers increasingly use AI.

Advertisement

“This incident highlights a rapidly emerging reality in software development: how AI coding assistants can inadvertently introduce workflow injection vulnerabilities, and how automated AI agents can rapidly surface them in the wild,” Nagli wrote.

Of course, the Google-owned biz has a vested interest in saying this. But this doesn’t make it not true.®

Source link

Advertisement
Continue Reading

Tech

One AI module faked 86% of a pipeline’s accuracy gains by feeding another the answers

Published

on

A retrieval-augmented generation (RAG) system is built to answer strictly from the documents it retrieves. But when engineers optimize these AI pipelines end-to-end, the reader module can learn a shortcut: instead of relying on retrieved evidence, it starts answering from its own internal memory — while the system’s overall accuracy keeps climbing. This is the hidden challenge of “role drift,” a failure mode in compound AI systems where individual modules learn to bypass their assigned tasks even as end-to-end performance improves.

To address this, researchers at MIT and Harvard introduce Role Anchor, a technique that forces modules to stay in their lanes during training. When applied, the technique mitigates role drift. For example, it forces the RAG reader to rely on retrieved evidence instead of answering based on its internal knowledge.

The primary takeaway for practitioners is that end-to-end accuracy alone can overstate how much a compound AI system has genuinely learned. Engineers must evaluate individual components and ensure they work as intended.

Role Anchor serves as both a guardrail and a diagnostic tool when optimizing multi-step LLM pipelines. It can be essential for real-world AI applications that require a strict division of labor between modules.

Advertisement

Why terminal accuracy hides the problem

Compound LLM systems divide complex tasks among specialized modules. For example, a system designed for multi-hop reasoning might split a task between a “Decomposer” and a “Solver.” The Decomposer breaks a large problem down into manageable sub-tasks, while the Solver computes the answers to those sub-questions. This division of labor allows AI engineers to delegate execution to smaller, cheaper models, and makes it possible to process sub-tasks in parallel where possible.

To improve the performance of AI pipelines, engineers typically optimize them using end-to-end reinforcement learning (RL) guided by a single “terminal reward.” This means the system is evaluated on whether or not the final answer is correct (the researchers call it “terminal accuracy”). When this terminal accuracy goes up, the system is considered to be learning and working as intended.

However, terminal accuracy does not verify whether the modules properly executed the tasks they were assigned. As Xiaoyang Cao, co-author of the paper, told VentureBeat, “Terminal accuracy reduces the behavior of an entire multi-part AI system to a single number. It shows whether the final answer is correct, but says little about which components contributed or whether they followed their assigned roles.”

This blind spot leads to role drift, a failure mode where a module’s behavior diverges from its assigned role during optimization, even though the system’s terminal accuracy continues to improve. 

Advertisement

“For engineering teams, the practical risk is that they can deploy a pipeline that passes every end-to-end evaluation even though its intended division of labor has silently broken down,” Cao said. Because the reward system only scores the final answer, it fails to detect or penalize the module for going rogue.

Role drift

Role drift (image credit: VentureBeat)

Consider how this happens in the Decomposer-Solver pipeline. The Decomposer’s assigned role is to write abstract sub-questions without solving the task, leaving the reasoning to the Solver. Under end-to-end RL, the Decomposer quickly learns that the weaker Solver is prone to errors on abstract tasks. To maximize the reward, the Decomposer begins leaking or planting answers into the sub-questions it sends to the Solver. The Solver ends up parroting the answer the Decomposer fed it. Terminal accuracy goes up, but the intended architecture is compromised.

But if the system is getting the right answers and accuracy is going up, why should we care if a module drifts from its role?

Advertisement

Real-world deployment requires much more than just a correct final answer on a training dataset. The implicit roles assigned to these modules ensure scalability, reliability, and auditability. Consider what happens when role drift takes over:

  • Loss of efficiency and auditability: In the reasoning example, role drift causes the Decomposer to do all the heavy lifting instead of planning and delegating. “Once the decomposer starts putting answers directly into its sub-questions, the solvers are reduced to copying those answers,” Cao said. “You are still paying to run [different modules], but they are no longer doing independent work.” The workload can no longer be parallelized across multiple Solvers, it cannot be delegated to cheaper models to save compute, and downstream human stakeholders can no longer audit the system’s logic step-by-step to verify how it arrived at the answer.

  • Fragility in dynamic environments: Consider a RAG system, in which a Reader model is tasked to answer questions strictly using external retrieved documents. If the Reader drifts and learns to rely on its own internal parametric memory instead (because its memory happens to be accurate during training), the system becomes brittle. When the enterprise updates its database with new information, or a user asks a question about a novel topic outside the model’s pretraining, the system will fail because it abandoned the grounding mechanism it was built to use.

How Role Anchor measures a role — and enforces it

“Training only for the final outcome rewards a system for producing the right answer, regardless of how it gets there,” Cao said. To counter this, Role Anchor serves as a lightweight regularization technique that makes role instructions part of the training objective. It compares how the component behaves with and without those instructions and discourages training from weakening their effect. 

At a high level, it ensures the module continues to respect the steering influence of its original role prompt throughout the reinforcement learning optimization process, making role drift both measurable and controllable.

A key insight of Role Anchor is that a role’s effect can be measured by comparing how a model behaves with and without the role prompt. The system evaluates two different prompts for each module:

Advertisement
  1. The specialized, instruction-heavy role prompt (e.g., “You are a careful Reader. Use the retrieved passages to answer the user’s questions…”).

  2. The neutral prompt (e.g., “Answer the user’s question…”).

For any given input, the model outputs a probability distribution for the next token. When run under the role prompt, it will favor certain tokens. When run under the neutral prompt, it behaves like a generic assistant. The difference between these two probability distributions is the “role utility.”

Role utility

Role utility (image credit: VentureBeat with Nano Banana Pro)

This utility measures the ”nudge,” or the direction and strength with which the role prompt shifts the LLM’s default predictions. If a token is highly aligned with the assigned role, the role prompt boosts its likelihood compared to the neutral baseline (or “nudges” the model toward that token).

Before starting RL training, Role Anchor keeps a frozen copy of the model as reference and measures the role prompt’s original nudge on this reference model. This pre-RL nudge serves as the ground truth of the designer’s intent, acting as a proxy for how the role prompt is supposed to steer the model.

Advertisement

During RL training, as the active model’s weights are updated, Role Anchor regularly calculates the current nudge and compares it to the reference nudge. If the current nudge starts to fade or deviate from the reference, Role Anchor applies a penalty to the model to prevent role drift.

Role Anchor

Role Anchor (image credit: VentureBeat with Nano Banana Pro)

To see this practically, consider the RAG system evaluated by the researchers. In this pipeline, the Reader module is explicitly instructed to answer user questions based only on retrieved documents, rather than relying on its internal knowledge.

During unconstrained, outcome-only RL, the reader learns that the upstream retriever is sometimes noisy. To maximize accuracy on the training set, it starts ignoring the retrieved passages and answering from memory. Consequently, the gap between its behavior under the role prompt and the neutral prompt shrinks to the point that the reader starts behaving identically under both, ignoring the grounding instructions.

Advertisement

In contrast, Role Anchor detects when the reader’s nudge deviates from the reference nudge. It applies a penalty, redirecting the model’s parameters away from this memory-based shortcut. This forces the reader to find role-compliant ways to improve, such as learning how to extract answers from the retrieved passages more robustly or avoiding using its internal knowledge when the retrieved passages are faulty.

The numbers: how much of the accuracy gain was real

To test the efficacy of Role Anchor, researchers evaluated it on the RAG and Decomposer-Solver (DEC) pipelines. The experiments compared systems trained with standard outcome-only reinforcement learning (no anchor) against systems trained with Role Anchor.

Under outcome-only RL, the RAG system’s terminal accuracy rose, but its internal integrity collapsed. The researchers measured “Evidence-Following Accuracy,” a probe testing if the model changes its answer when the retrieved text is deliberately swapped to state the opposite. This metric plummeted from 0.86 to 0.54 (just above random chance), meaning the model learned to ignore retrieved passages and rely on its pre-trained parametric memory instead. In one test, researchers deliberately changed a piece of information in a retrieved document to contradict the model’s internal knowledge. The unanchored model did not update the response because it wasn’t using the external document.

When Role Anchor was applied, the Reader’s Evidence-Following Accuracy remained at 0.869, proving it relied strictly on the retrieved text. When researchers fed the anchored model random passages that were unrelated to the input prompt, its accuracy correctly dropped because it refused to use its internal knowledge. The unanchored model scored higher on random passages because it was guessing from memory.

Advertisement

The Decomposer (DEC) pipeline showed an even more dramatic failure mode. Under outcome-only RL, terminal accuracy shot up, but the “insertion rate” (i.e., the frequency at which the Decomposer leaked the answer into the sub-questions it sent to the Solver) surged from 0.143 to 0.596.

Role Anchor performance

Role Anchor makes sure the model stays in its lane throughout RL training (source: arXiv)

In the RAG pipeline, preserving the intended role cost the system a very modest accuracy drop (-0.067). The Reader still learned to be better at extracting answers, but it did so legitimately rather than by cheating with its internal memory. This means it is more reliable on real-world tasks with novel knowledge it has not seen during training.

In the DEC pipeline, unanchored RL improved accuracy by 0.310 above the base model, while Role Anchor only showed a 0.057 improvement. When diagnosed, it turned out that the underlying issue was that the Solver model was too small and couldn’t learn the problem-solving part. This forced the Decomposer model to cheat and provide the answer to boost the terminal accuracy. This meant 86% of the unanchored improvement was fake, and the system had simply learned to exploit a shortcut instead of learning how to reason or decompose problems better.

Advertisement

However, this tradeoff is not a universal rule. In some cases, eliminating shortcuts can actually boost overall performance. “Role Anchor… does not necessarily reduce final accuracy,” Cao said. “In a coding pipeline we recently tested, the model had learned to manipulate its own test executor during reinforcement learning training. Adding Role Anchor completely eliminated that shortcut while slightly improving correctness on the final tests used to judge the code.”

What it takes to add Role Anchor to an existing pipeline

For engineering teams looking to apply this technique, “Role Anchor can be added to an existing reinforcement learning fine-tuning process as an extra training objective for each component that a team wants to anchor,” Cao said. The main pipeline and deployment setup remain entirely unchanged.

To implement it, engineers need three specific items for each anchored component: its original role instructions, a matched neutral version with the role information removed, and a saved copy of the model from before reinforcement learning fine-tuning.

Importantly, there is no latency penalty at inference time. “Role Anchor runs only while the model is being trained, so it does not slow down the deployed system,” Cao said. He noted that their current implementation takes roughly 20 percent longer during training due to additional calculations, though there is likely room to optimize and reduce that overhead. The research code, training configurations, and selected model weights will be released publicly in the near future.

Advertisement

Deciding when to use Role Anchor is a case-by-case decision based on whether final accuracy captures everything that matters. Cao points to a regulated legal RAG system as a prime candidate. “The component producing the answer may need to follow retrieved evidence, stay grounded in an approved set of documents, and produce answers that can be traced back to their sources,” he said. “Final accuracy alone cannot verify those properties, so the behavior of that component needs to be measured and enforced directly.”

As enterprise AI evolves toward more complex compound pipelines, role enforcement will become harder, and relying on prompts alone will prove unreliable. “At larger scales, role specifications will need to be enforced through both training and system design,” Cao said. “Methods such as Role Anchor can help preserve intended behavior during training, while clear system boundaries, limited tool permissions, and monitoring during use can provide additional safeguards.”

Source link

Advertisement
Continue Reading

Tech

US Grid Operator PJM Proposes Forcing Data Center Off Grid During Emergencies

Published

on

An anonymous reader quotes a report from Reuters: PJM Interconnection, the biggest U.S. grid operator, proposed on Thursday a new framework that would force data centers to use back-up generators when electricity supply on the grid approaches dangerously low levels. The grid operator’s proposal dovetails with President Donald Trump’s Ratepayer Protection Pledge, a non-binding initiative to protect residential customers from getting saddled with costs related to data center power consumption, PJM said.

A new emergency procedure would notify utilities to reduce or transfer the electricity demand from data centers and other large power users ahead of any action that would shut off traditional consumers such as households. PJM said it does not, however, currently have the authority to curtail power to those sites and would require the cooperation of individual state governments.

PJM manages the electricity for 67 million people in a territory that stretches from Washington, D.C. to Chicago. Its proposal highlights a growing tension between the rapid expansion of data centers and the ability of the nation’s power grid to keep up. If PJM cannot close its supply gap, millions of residents and businesses face an increased risk of blackouts, and the cost of new generation could be passed on to other power consumers. At its recent capacity auction, PJM hit its $325-per-megawatt-day price cap but still came up about 6.8 GW short of its projected reliability needs.

With rapidly expanding data centers adding pressure to the grid, PJM has also proposed creating a registry to track their locations and power consumption.

Advertisement

Read more of this story at Slashdot.

Source link

Advertisement
Continue Reading

Tech

The US is trying to kick foreign robots out, but the local supply chain might not be there yet

Published

on

The FCC has a track record of using its Covered List to squeeze Chinese tech out of American markets, and robots just became its newest target. Foreign-made humanoids, quadrupeds, and even robot vacuums now need to be mostly built in the U.S. to sell here (via Rest of World). 

So what does the new rule actually require?

The FCC added “advanced robotic devices” to its Covered List last month, the same national security roster that’s previously blocked devices from Chinese companies like Huawei and ZTE. New models of foreign-made humanoids, quadrupeds, robot vacuums, and even lawn mowers are now barred from entering the U.S. market.

Only models assembled domestically and sourcing at least 65% of their component value from within the U.S. qualify. The threshold will climb to 75% by 2029. The policy doesn’t target any single country (or mention one), but it’s widely understood as aimed squarely at reducing Chinese imports.

The good news is that existing products (already on sale) aren’t affected. Robots imported purely for research and development (not for commercial sale) stay exempt from the rule for now.

How are startups actually reacting to this?

Founders say hitting 65% local sourcing threshold is nearly impossible right now. Plenty of motors, sensors, and actuators are either too expensive, too slow to source, or, the worst one of all, simply unavailable domestically. The issue is so severe that employees have reportedly flown parts over from China in their luggage just to keep prototypes moving. 

Advertisement

Michael Perry of Persona AI put it bluntly: “You need to provide the carrot as well as the stick.” Not everyone’s upset, though. Oregon-based Agility Robotics and San Francisco’s Nori Robotics both welcome the rule, stating that it will eventually shield them from cheaper Chinese competition. 

However, there’s no denying that the rule complicates sourcing for the immediate future. China didn’t end up controlling robotics manufacturing by accident. Decades of state investment, a deep engineering workforce, and a boom in electric-vehicle production carved out the kind of sensor, battery, and actuator supply chain robots also rely on. 

The result: Chinese-made humanoids account for nearly 90% of everything sold worldwide in 2025, according to research firm Omdia. 

Source link

Advertisement
Continue Reading

Trending

Copyright © 2025