Connect with us

Tech

How To Get The Best Battery Life From Your Retro Handheld Game Console

Published

on

If your vintage gaming system eats through batteries in a hurry, this is for you.

Old-school gaming handhelds can be battery hogs. Whether it’s the original Game Boy’s four AA batteries or the Game Gear’s six, they tear through cells at a rate that makes modern consoles feel like luxuries. Here’s the easiest way to boost your battery setup on your old Nintendo, Sega or Atari portable.

For this guide, we’re focusing on the most popular battery-powered handhelds from the late ’80s through the early 2000s. Those would be Nintendo’s original Game Boy, Game Boy Color, Game Boy Pocket and Game Boy Advance, plus Sega’s Game Gear and the Atari Lynx.

Advertisement

The bottom line? Modern rechargeable batteries smooth much of the friction out of these retro handhelds. They save money, cut environmental waste and let you avoid the headaches that go along with disposable alkalines.

The right rechargeable batteries

For most people, a set of nickel-metal hydride (NiMH) rechargeable batteries is the way to go. They’re reliable, reusable and much cheaper over time than disposable alkalines.

Panasonic Eneloop is an Engadget staff favorite. It’s consistent and tends to hold up over time better than cheaper options. If you want to save a few bucks, IKEA LADDA is a good budget alternative. (Some have speculated that they’re rebranded Eneloops, but that isn’t confirmed.) Both are sold in the AA and AAA sizes you need for these systems.

Advertisement
  • Game Boy – 4 × AA
  • Game Boy Color – 2 × AA
  • Game Boy Pocket – 2 × AAA
  • Game Boy Advance – 2 × AA
  • Game Gear – 6 × AA
  • Atari Lynx – 6 × AA

Just keep in mind that the system’s battery indicators don’t always behave the same way with rechargeables. If the low-battery warning comes on, you may want to get in the habit of quickly saving your game to be safe.

As for how many batteries you need, a good rule of thumb is to buy two full sets. That way, you always have one in the handheld with another charged and ready to go.

One possible exception is modded retro handhelds, since things like brighter screens can run through batteries faster than the stock hardware. In that case, a third set could be in order. Some power-hungry or voltage-sensitive mods might work better with 1.5V rechargeable lithium batteries, although they tend to cost more than standard NiMH cells.

Advertisement

USB-C battery mods

Some models, including the original Game Boy, can be modded with USB-C rechargeable battery pack kits. These make the handheld feel more like a modern system, removing the annoyance of battery swaps. Some work with USB-C power banks, although not all support charge-and-play.

There are a few risks and downsides. For example, some USB-C battery mods don’t fully isolate the battery while charging. This can generate extra heat and make it less safe to play while charging. So, it’s worth checking user feedback and sticking with well-regarded mods if you go that route.

Source link

Advertisement
Continue Reading
Click to comment

You must be logged in to post a comment Login

Leave a Reply

Tech

Security researchers scanned the Polish web and found courts, hospitals, and airports at risk of hacks

Published

on

Two Polish security researchers wanted to find out how vulnerable their country’s internet was to potential cyberattacks and quickly found that thousands of public agencies and websites were at risk of being hacked.

At the Def Con cybersecurity conference in Las Vegas on Friday, security researchers Robert Kruczek and Kamil Szczurowski said they wanted to understand the state of Poland’s public web out of a sense of patriotism and a desire to make it safer for everyone. 

Before long, the duo discovered more than 10,000 affected public entities with 250,000 websites with security flaws, including airports, hospitals, and government offices. 

The researchers found that some of the vendors’ buggy software, coupled with a lack of bug bounties and ways to report security flaws, are putting Poland’s public services at risk of hijacks and other attacks. The researchers also said that some bugs were incredibly easy to exploit but were not always taken seriously, with some vendors describing the bug reports as inconveniences.

Advertisement

The research comes as Poland is trying to shore up its cyber defenses after a wave of suspected Russian hacks targeting the country’s energy and water providers. Some of the hacks have been carried out by taking advantage of weak cybersecurity. 

Kruczek and Szczurowski found critical vulnerabilities in the widely used content management system Pad CMS, which allowed them to easily access over 300 public websites without needing a password. The software developer did not patch the software because it had become “end of life” and was no longer supported.

Another bug allowed them to gain access to the websites of some two-thirds of Poland’s judiciary, or about 245 courts, they said.

The duo reported their findings to the government through various official channels.

Advertisement

The researchers said during their talk that it was ultimately worth the hassle, saying that as a result we are “a little bit more safe.”

When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.

Source link

Advertisement
Continue Reading

Tech

Should I really believe the fitness age my smartwatch has given me?

Published

on

Smartwatches and fitness trackers are generally pretty great, as they’re packed with plenty of features that promise to offer a solid overview on your health.

Among the likes of heart rate monitoring and workout tracking, many of the best smartwatches and best fitness trackers provide users with a so-called fitness age which has the power to either give you bragging rights or keep you incredibly humble.

Having said that, how much weight does a device’s calculated fitness age really have? Should you take the age with a pinch of salt, or is there factual evidence behind the number on the screen?

I’ve explored the science behind the so-called fitness age and found out whether you really need to worry about it. 

Advertisement

How does a smartwatch calculate fitness age?

The exact measurements that go into a smartwatch’s fitness age calculation can vary slightly from device to device. However, calculating fitness age ultimately comes down to your VO2 max measurement, which is the maximum volume of oxygen that your body can use per minute. The higher the figure the better, as this means your body is generally able to take in more oxygen during more strenuous workouts.

Advertisement

With this in mind, it’s unsurprising that Garmin assesses your VO2 max figure to accurately decipher your fitness age. Garmin explains it does this by comparing your figure to those of people “of different ages within your same gender”. 

However, Garmin states that on its newer devices, fitness age takes additional factors into account. These factors include activity intensity, your resting heart rate and, controversially, your BMI.

Advertisement

The latter is controversial, as many have criticised BMI as being an outdated way of measuring someone’s actual health. In fact, a 2023 study revealed that BMI “does not properly assess body fat percentage and muscle mass or distinguish abdominal fat from gluteofemoral fat” and explains that women and minorities were either “not included at all, or only comprised a small percentage of the 7462 men” within a 1972 study which sparked the modern use of BMI in research and medical settings. 

Whoop app Healthspan ageWhoop app Healthspan age
Whoop age. Image Credit (Trusted Reviews)

Alongside VO2 Max levels, Whoop analyses your resting heart rate, sleep pattern and the amount of time each week you’ve spent exercising to determine your Whoop age. In addition, Whoop also factors in your lean body mass, if you’ve entered that detail into the app. 

Otherwise, devices like Apple Watch and Fitbit don’t actually offer a built-in fitness age. Instead, you can download third-party apps for your Apple Watch which measure your potential fitness age, while Fitbit offers a cardiovascular fitness score rather than an age. 

Advertisement

What is VO2 max, and do wearables provide an accurate measurement?

So, although smartwatches assess multiple factors, your individual VO2 max level is one of the key measurements to determine your fitness age – but how do devices actually measure this?

Advertisement

Both Garmin and Whoop disclaim that their VO2 max level is an estimate and isn’t totally accurate. Even so, both do provide insights into how they reach their respective figures. Whoop, for example, explains it estimates VO2 max by using its own three-tier proprietary algorithm that analyses your health and fitness data and “your demographic data including your age and biological sex”. Whoop promises that its VO2 Max tracking is within 3.3 to 3.7mL/kg/min of a lab-measured value. That roughly translates to around 7 to 8% of error.

Garmin, on the other hand, explains that its VO2 max estimations rely on an accurate user profile and your elevated heart rate. While Garmin doesn’t publicly reveal its level of accuracy, I’ve seen various reports of anywhere between 5 to 15% of error. However, a study conducted back in 2024 determined the accuracy of VO2 max estimates specifically with the Garmin Forerunner 265, which suggested that Garmin “may overestimate oxygen consumption compared to values acquired in the laboratory”.

Garmin VO2 max

Whoop VO2 max

Advertisement

I have a Whoop MG and a Garmin Cirqa, so had a look to see how each device has determined my VO2 max level. Interestingly, they differ slightly with their exact measures, with Whoop claiming my VO2 max is 46 while the Cirqa states a result of 43. Though Whoop’s estimate is higher, my fitness age is slightly older than Garmin’s at a very precise 28.9 years old. Garmin, ever the flatterer, says I have the fitness of a 26 year old.

Although the results aren’t too dissimilar from one another, the fact they vary slightly confirms that VO2 max and fitness age shouldn’t be taken as gospel, as two devices I wear constantly have provided me with different results.

Advertisement

Advertisement

Verdict – does fitness age matter?

Considering VO2 max is one of the key factors in determining an individual’s long-term health, the fact that Whoop and many of the best Garmin watches have the ability to estimate your level is undeniably brilliant. Sure, you can’t expect a 100% accurate reading, but an approach is surely better than nothing, right?

However, that error margin should be factored in when receiving your estimate, and therefore your fitness age, as it means the levels are likely not completely accurate. They might not be too far off, with Whoop promising up to 8% of error, but that’s still enough doubt that you shouldn’t completely rely on the ages provided. Not only that, but remember devices factor in additional measurements that are seen as outdated in modern terms, like BMI.

Essentially, we’d advise that while your fitness age is a solid indicator of your overall fitness level, it should be taken with a pinch of salt. That’s not to say you should ignore an age that’s double your actual age, nor that you shouldn’t be pleased with one that’s younger, but just keep in mind that it isn’t actually a completely reliable measurement.

Advertisement

Source link

Continue Reading

Tech

US Navy’s drone boats tracked an $81 million cocaine shipment before warships swept in for the Caribbean interception mission

Published

on


  • Voyager drone boats tracked smugglers before US Navy forces intercepted the shipment
  • Autonomous vessels helped seize cocaine worth more than $81 million offshore
  • Electric-powered drone boats can patrol continuously for more than 100 days

Unmanned surface vessels have helped the US Navy intercept a cocaine shipment worth more than $81 million in the Caribbean Sea.

Two Voyager drone boats, built by California-based Saildrone, tracked a vessel of interest before coordinating directly with naval forces stationed nearby.

Source link

Advertisement
Continue Reading

Tech

What’s The Minimum Distance You Should Leave Between A Washing Machine And The Wall?

Published

on





The right washing machine is all about space. Enough space on the inside, first and foremost, but it’s just as much about the amount of available space in your laundry room, too. That’s the part that determines whether the appliance even fits properly, not to mention operates safely. It’s one of those mistakes everyone makes when installing appliances: Too many homeowners focus only on the machine’s internal space without thinking about this second part at all. As a rule of thumb, you should always consider the external space, too.

Without the right amount of clearance behind the washer, you won’t have enough room for water hoses, power cords, or proper ventilation. On installation, you should leave at least 6 inches of space between the back of a washing machine and the wall. That gap leaves just enough space to connect water lines and electrical cords without crimping or damaging them. You also need to leave that room for airflow and whatever vibration might come from the machine during operation.

Advertisement

Other measurements to know when installing a washing machine

In addition to those 6 inches of space in the back, you should also leave about an inch of clearance on each side of the washer. It needs that space for ventilation, most importantly, but it also helps with noise reduction and damage control. Without it, the machine might bump up against the dryer or the wall… especially if yours tends to walk across the floor

Because there’s no universal standard washer size across the major washing machine brands, dimensions are going vary tremendously depending on whether it’s a front-load, top-load, compact, stackable, or large capacity machine. That’s why it’s essential to take measurements before buying and installing. Of course, be sure to factor the 6-inch and 1-inch guidances into your measurements.

Don’t forget about door clearance, either. It’s easy to forget, but the machine is going to have to both fit through the door, and also have enough space around the machine for the door to swing open all the way. You’ll need at least 20 inches of open space in front of or above the appliance as well as an unobstructed path through whatever doorways, hallways, or tight corners you’ll need to get through upon delivery.

Advertisement



Source link

Advertisement
Continue Reading

Tech

NHS Tayside investigates breach concerning data of dead girl

Published

on

security

Investigation into whether staff improperly accessed Minnie Merriman’s file after she was named for the first time this week

A Scottish NHS trust is investigating a data breach concerning the medical records of a nine-year-old girl who died earlier this week and was named publicly for the first time on Wednesday after a man was charged with her death.

The alleged breach occurred at Ninewells Hospital in Dundee, and reportedly involved staff members accessing the girl’s medical records without authorization or clinical need.

Advertisement

A spokesperson for NHS Tayside, which oversees Ninewells Hospital, said: “NHS Tayside is currently investigating the circumstances of an alleged data breach which happened in a working clinical area where staff access patient information.

“As a matter of governance, any data protection breach would be recorded and investigated by NHS Tayside and, where appropriate, reported to the Information Commissioner’s Office (ICO). It would not be appropriate for us to comment further on individual staffing matters.”

NHS Tayside did not respond to questions about the nature of the accessed data nor who is thought to be behind the intrusion.

Medical records in the UK are protected by the UK GDPR, contained in the Data Protection Act 2018 as well as several common law confidentiality rules. NHS staff are only allowed to access patient information where there is a legitimate clinical or other work-related need.

Advertisement

A 35-year-old man whom police say was known to the child, was arrested and appeared in court on August 5 over the death of Minnie Merriman. 

The man issued no plea at Forfar Sheriff Court on the day of his arrest and has been remanded in custody.

Merriman was found in Elliot Industrial Estate at approximately 0002 on Monday, August 3, with serious injuries.

The young girl was then taken to Ninewells Hospital in Dundee, where she later died.

Advertisement

Police Scotland said that they are not currently looking for anyone else in connection with her death.

Other members of Merriman’s family, who are from West Yorkshire and were camping nearby, are being supported by specialists. 

A family statement, released through Police Scotland, read: “We are devastated with the loss of our beloved, absolutely incredible, beautiful and brave Minnie Moo. Our family asks that our privacy is respected at this extremely difficult time.”

Detective Inspector Mike Ness of Police Scotland’s major investigation team said: “Our thoughts remain with everyone affected by these events, especially Minnie’s family.

Advertisement

“A police presence will remain in the area while our enquiries continue.

“Anyone with any concerns, or information, should approach these officers or contact Police Scotland on 101, quoting incident number 0008 of Monday, 3 August 2026.” ®

Source link

Advertisement
Continue Reading

Tech

Apple’s legal battles with Jon Prosser encounter new hurdles

Published

on

Jon Prosser was sued by Apple over the alleged theft of pre-release information, saw a default ruling due to inaction, had that overturned, and is once again failing to provide discovery materials. There’s a reason this time.

Life hasn’t been easy for Jon Prosser since a July 2025 lawsuit from Apple accused him and Michael Ramacciotti of stealing information from Apple employee Ethan Lipnik’s test device. He says he’s not guilty, but has created roadblocks at every possible point in this case.

According to a new filing, Jon Prosser has not been responding to requests for discovery since he rejoined the case in June. Apple last heard from Prosser’s counsel on July 6, 2026.

In the interim, Prosser had a second child. The lack of response is attributed to the newborn.

Advertisement

He allegedly hasn’t had time to provide requested materials. In the filing he says that he “will find dates when additional discovery may be provided to Apple.”

However, given the situation involves a newborn, Apple appears to be a little more understanding of this delay. The joint filing suggests that Apple has approved the delays in discovery and the new dates.

What isn’t clear is whether Apple will take a softer approach to the final determination and punishment for Prosser. The worst case here is that he is barred from reporting on Apple, which is a significant portion of his business.

A new deposition is set to take place in September 2026. The next legal filing is expected on October 7, 2026.

Advertisement

Source link

Continue Reading

Tech

TP-Link Roam 7 travel router review

Published

on

Why you can trust TechRadar


We spend hours testing every product or service we review, so you can be sure you’re buying the best. Find out more about how we test.

TP-Link Roam 7: 30-second review

The TP-Link Roam 7 is a compact travel router that lets you connect to whatever internet you can find and turns that into your own private network. This helps reduce the need to reconnect all your devices in each new location; just connect the router and all other devices will connect automatically through it.

The Roam 7 differs from a mobile router insofar as there’s no internal SIM or battery; power is supplied through a USB source such as a power bank, and the internet connection comes from a range of available options, including Ethernet at a hotel or business, public Wi-Fi, your phone tethered or a USB modem.

Advertisement

Source link

Continue Reading

Tech

Cloudflare says humans could become a “rounding error” as bots generate 1,000 times more internet traffic

Published

on

Crystal ball: Just a couple of months after Cloudflare revealed that bots have officially surpassed human traffic online, the company has given its view on what things will look like in five years. It’s not good news for us fleshy meatsacks: non-human traffic could be up to 1,000 times higher than human traffic.

CDN and cybersecurity giant Cloudflare held its Q2 earnings call on Thursday, during which chief financial officer Thomas Seifert gave a concerning prediction about what the web will look like in the future.

“If the current trends continue, we think in five years, non-human traffic will be as much as 1,000 times as much as human traffic,” he said.

“In other words, humans will be a rounding error on the internet, not because human traffic goes down, but that’s just how fast we’re seeing non-human traffic grow.”

Advertisement

Seifert did add the important caveat that his predictions have been wrong in the past. Cloudflare CEO and co-founder Matthew Prince will likely be able to relate to this. He had predicted that bot traffic would surpass humans at the end of 2027, but it happened earlier this year.

Also read: The Zero Click Internet

It goes without saying that the sudden acceleration in non-human traffic is being driven by AI – agentic AI, specifically. Agentic systems behave very differently from traditional web crawlers and fraud bots. Their activity can closely resemble ordinary browsing, but it happens at machine speed and can be repeated on an enormous scale.

Advertisement

Even a straightforward prompt may trigger thousands of requests. Prince illustrated the difference using online shopping: someone searching for a camera might check five retailers, while an AI agent could examine 5,000 sites on their behalf. Multiply that behavior across millions of users asking assistants to compare products and airfares, research topics, collect information, or condense articles, and you can see why Seifert believes humanity’s online presence will become a rounding error.

Seifert said that if traffic does grow the way he expects, “we’ve got to get a lot more efficient,” which is where Cloudflare’s services come in, of course. He also mentioned the security implications of a web made up pretty much entirely of bots.

The idea that Dead Internet Theory – the concept that the internet is mostly run by bots and AI – will be demonstrably true in a few years (or now, arguably) is a concerning one. For a start, it could see the end of independent sites that have relied on the web’s business model since the 1990s.

There’s also the fact that an internet full of bots, AI slop, and privacy risks is making more people pull away from the online world. In our poll, 97% of people said the internet today was a hellscape compared to its early days, and it looks like things are only going to get worse.

Advertisement

Source link

Continue Reading

Tech

Google Should Still Be Forced To Shed Chrome, Advocacy Group Argues

Published

on

“Google should be required to divest the Chrome browser, and prohibited from paying Apple to distribute Google’s search engine, the nonprofit advocacy group Public Knowledge argues in a new court filing,” MediaPost reports, citing a friend-of-the-court brief filed Tuesday in the D.C. Circuit Court of Appeals:

The group adds that “independent ownership” of Chrome “would open the distribution channel Google controls and allow Chrome to serve browser users when it makes privacy decisions and determines how to integrate search and (artificial intelligence)…”

In September 2025, [U.S. District Court Judge] Mehta issued a remedies order that requires Google to share some data about users’ searches with “qualified” competitors and to provide syndicated search results and ads to those competitors. The order also prohibits Google from entering into exclusive distribution contracts for Google Search, Chrome, Google Assistant and the Gemini app for six years, but allows the company to continue to make payments for search-ad revenue or distribution to Apple, Mozilla and others…
Google recently appealed Mehta’s order. The company argued in its written brief that it “prevailed in the marketplace fair and square,” adding that Apple and Mozilla “sensibly chose” Google as the default search engine “because it gave their users the best experience,” and because Apple and Mozilla would earn the most ad revenue through the deals. The [U.S.] Justice Department and states countered to the appellate court last week that the liability finding should stand, and also argued that Google should have been banned from paying Apple and Mozilla for placement as the default search engine on their browsers.

[Antitrust enforcers had originally asked the judge to order Google to divest Chrome, but he’s already rejected that request.] The government did not argue in its appellate papers that Google should be forced to sell Chrome. But Public Knowledge independently contends in its friend-of-the-court brief that divestiture would benefit consumers… “Divestiture would place those decisions with an institution whose success depends on serving browser users primarily….” The group is calling the appellate court’s attention to Google’s April 2025 decision to preserve tracking cookies — a reversal from its earlier plans to block third-party cookies by default. “Google is in the position of both deciding Chrome’s tracking rules while running the advertising business affected by them,” Public Knowledge writes. “An independent Chrome could make those decisions on behalf of users alone.”

Advertisement

But Firefox developer Mozilla filed its own friend-of-the-court brief Thursday warning Firefox could be forced to “exit the browser and browser engine markets” if it can’t receive payment from Google for distributing its search engine, according to a later report from MediaPost:

Federal and state antitrust enforcers recently asked the appellate court to reverse the portion of Mehta’s order that allows those payments to continue. But Mozilla counters in its new friend-of-the-court brief that Mehta’s decision regarding those payments was supported by the evidence –including a study it conducted concluding that its revenue would “decline dramatically” if forced to replace Google with Bing as Firefox’s default search engine…

Google is expected to file new arguments with the appellate court next month.

Read more of this story at Slashdot.

Advertisement

Source link

Continue Reading

Tech

Qwen 3.8-Max and Claude Opus 5 show why raw benchmark scores don’t predict the bill

Published

on

Alibaba released Qwen 3.8-Max this week and marketed the preview as second only to Claude Fable 5 (their launch-day table was more equivocal: the model leads on one of 12 coding-agent rows). But an independent harness came close to the opposite conclusion: a benchmark run, apparently using the Preview version, put Qwen 3.8-Max’s best effort setting mid-pack, and its default setting last.

Both results are real and defensible. The gap between them is about token and time budgets, and that matters because those figures aren’t usually headline numbers. Alibaba’s footnotes give its coding numbers a five-hour timeout, and up to 12 hours per run on PaperBench. The independent harness, VulcanBench, allowed between 45 and 60 minutes of wall clock time. A time budget between five and 16 times larger on Alibaba’s side explains the huge difference in results.

It’s time to do two things to start accounting for these differences when choosing models. First, the metric to use is cost per successful task: total spend, including everything you spent on attempts that failed, divided by the tasks that actually passed your acceptance check. Second, you need to make time or token budgets an explicit part of your acceptance criteria, not a hidden detail.

Price per token has stopped predicting the bill

The comparison everyone published in Qwen 3.8-Max’s first week was a price comparison, because that was the only data available. It is not a cheap model. DeepSeek-V4-Flash-0731, which entered public API beta on July 31, lists at 14 cents per million input tokens and 28 cents output. Qwen 3.8-Max lists at $2 and $6. Kimi K3 sits at $3 and $15.

Advertisement

Those prices tell you less than they used to, for a reason specific to reasoning models like Qwen: getting to a result costs thinking tokens. A model that spends most of its token allowance on reasoning can reach a token cap before it writes the answer, giving you an empty result indistinguishable from a total failure at the cost of a full run.

Artificial Analysis has the cleanest published measurement of how this can affect real agent spend: running its Intelligence Index on DeepSeek-V4-Flash at maximum effort took 210 million output tokens against a class median of 100 million. Absolute cost stayed low anyway, because the tokens were so cheap. But verbosity costs time, not just money, and depending on your use case that can sink you.

What you need is a number that counts everything you spent, including the attempts that came back empty, against the tasks that actually got done in the time and token budget you specified. This is what a cost-per-success metric helps you see.

Your failure rate is partly a configuration setting

A run that produces a wrong answer and a run that runs out of budget are different events with different fixes. Almost no harness distinguishes them, and almost no leaderboard reports the split. I hit this building an agent benchmark of my own: the harness logged a failure and nothing about why, and I had to add the distinction myself. When you do separate them, budget exhaustion turns out to dominate.

Advertisement

Long-Horizon-Terminal-Bench, published in July, ran 17 frontier models across 46 tasks through a shared harness with one 90-minute attempt each. Timeouts accounted for 79% of unresolved runs, against 19% for agents that stopped on their own and 3% for harness errors. The authors are careful about what that does and does not mean: the timed-out runs were not close to finishing, with mean reward between 0.10 and 0.35, so you cannot assume more time would have resulted in success. But the lesson is: benchmarks are implicitly measuring time efficiency, whether or not they shout about that.

The clearest published example of the mechanism comes from VulcanBench, the same open-source harness behind the Qwen chart. In a report dated July 26, Claude Opus 5’s lowest-effort setting was its best, solving 20 of 23 tasks against 18 at high effort. The extra reasoning wasn’t useless: high effort returned the fewest wrong answers of any setting, one against three. It ran out of clock instead, and a timeout scores zero. Two of its three regressions were cutoffs on tasks that low effort solves, and given unlimited time on both it only ties its cheapest setting, at 3.1 times the cost.

That has a direct consequence for anyone building a routing ladder. The standard design escalates to more reasoning when a cheap attempt fails, on the assumption that the next rung is better and merely costs more. For a meaningful share of model and task combinations that assumption is wrong, and you pay the higher rung’s price to escalate into a timeout or hitting a cap.

Who is already measuring this

Several groups have landed on cost per successful task independently in the last few months, which is the strongest signal it’s becoming standard.

Advertisement

VulcanBench reports dollars per solved task as a headline column and has since its earliest reports. Long-Horizon-Terminal-Bench publishes per-task cost next to accuracy, and its most instructive row is GPT-5.4 at roughly $26 per task with a much lower pass rate than Grok 4.5 at about $11. TestEvo-Bench runs agents under a cost cap, and Claude Code’s test-generation score falls from 71% to 44% at the tighter cap.

Vendors are already on board with the idea of measuring per successful task. HubSpot moved its Breeze Customer Agent in April to 50 cents per resolved conversation, down from $1 per handled conversation. Zendesk bills per automated resolution. Fin charges 99 cents per outcome and bills only on end-to-end resolution.

What to change this week

  • Emit a failure reason on every agent run as a required field, with budget exhaustion, verifier failure and harness error as distinct values rather than one failure flag. Until you can separate a timeout from a wrong answer, your pass rate is measuring two things at once and you cannot tell which one to fix.

  • Compute cost per successful task per effort level, not just per model. Total spend including failed attempts, divided by tasks that passed your acceptance check. The ranking will not match the rate card, and the cheapest setting may well win.

  • Cap on tokens rather than wall clock unless latency is genuinely in your service level objective. A wall-clock cap scores your provider’s serving speed as model quality.

  • Check the default effort setting on everything you have deployed. Qwen 3.8-Max runs at its highest reasoning setting when the effort field is unset, and its highest setting was its worst performer in independent testing. A team that never touches that parameter is running the configuration that costs the most per solved task.

Source link

Advertisement
Continue Reading

Trending

Copyright © 2025