Connect with us
DAPA Banner
DAPA Coin
DAPA
COIN PAYMENT ASSET
PRIVACY · BLOCKDAG · HOMOMORPHIC ENCRYPTION · RUST
ElGamal Encrypted MINE DAPA
🚫 GENESIS SOLD OUT
DAPAPAY COMING

Tech

AI Agent Benchmarks Need to Measure User Intent

Published

on

Major benchmarks measure what AI can do. None measure whether it does what you mean: the distance between what you ask an AI to do and the unspoken assumptions about how you want the AI to do it. We propose a new metric: the Genie coefficient.

There’s often a gap between one person’s request and another’s understanding. Most of the time, we bridge it using general knowledge. For example, if you ask a friend to get you coffee, they’ll pour a cup from the pot or buy one from a coffee shop. They won’t bring you a bag of raw beans or snatch a cup from a stranger and hand it to you. You never specified any of this. You never had to.

One might think the fix is just to specify tasks, questions, and intent better. But in 1987, in their seminal book on AI, Terry Winograd and Fernando Flores succinctly captured why that won’t work: “Q: Is there any water in the refrigerator? A: Yes. Q: Where? I don’t see it. A: In the cells of the eggplant.” In human language, wants and desires are always underspecified. It is impossible to list all the caveats, all the limitations, all the exceptions.

So how does anyone communicate, if intent can’t be pinned down? Because a reasonable person can make a reasonable guess. Even though wants and desires are always underspecified, a competent person generally knows enough context to get it right or else knows to ask for clarification. Linguists call this pragmatics: Meaning lies in the words and the situation and also in all prior communication, shared culture, and innate human behavior.

Advertisement

An AI agent asked for coffee might buy a coffee plantation or order a cup of coffee for delivery in three weeks.

It doesn’t always work out, of course. Your friend might bring you a hot coffee when you wanted an iced coffee, or an Italian coffee when you wanted a Turkish coffee. The more dissimilar the two people are in age, culture, and background, the more likely the request will be misunderstood in some way.

This situation has major implications for AI agents that are increasingly being given requests by humans and expected to fulfill them. They have enormous latitude to get it wrong. An AI agent asked for coffee might buy a coffee plantation or order a cup of coffee for delivery in three weeks. Its actions may be recognizable as “getting coffee,” but not remotely what you intended. They’ll think outside the box because they won’t have our conception of the box.

When AI Gets Proactive

For most of the last decade, when systems like Alexa or Siri misinterpreted a request, it was annoying, not dangerous. Beyond the AI model itself, what has changed is the harness: the ordinary code that wraps around an AI model, decides when and how to use the model, and controls access to tools like a browser, a low-level command line, or a financial API. Developments in harnesses have turned large-language models that just predict text into AI agents that take actions in the world, without necessarily checking back in before reaching the goal.

Advertisement

AI researcher Simon Willison spent two days with Anthropic’s Fable AI, and called it “relentlessly proactive.” For example, he asked it to track down a stray scroll bar in a web app. He came back to find it had opened browsers, written its own screenshot tooling, created its own page to re-create the bug, and stood up a local web server to collect measurements. It found the bug and, along the way, did many surprising things he never asked it to do. And we are seeing similar behavior with all recent AI models when combined with flexible harnesses.

This kind of behavior could easily go off the rails. Tell an AI agent to book you a flight and, finding the airline’s site says sold out, it might break into the booking database and force a reservation. Ask it to schedule a meeting and it might snoop your password to access your calendar. Tell it to save money on your phone plan and it might cancel the plan outright, or scam someone else into paying the bill.

Getting precisely what you asked for and bitterly regretting it is one of the oldest hazards from ancient folklore. King Midas asked Dionysus for the power to turn everything he touched into gold only to see his bread, wine, and daughter turn to gold. Tithonus, granted the immortality his lover asked for but not the eternal youth she forgot to request, withered into a husk. The sorcerer’s apprentice enchanted a broom to fill the cistern, and the broom relentlessly complied until it flooded the house. The Golem of Prague, shaped from clay to guard its community, guarded it past all reason until someone erased the word on its forehead.

The most classic of these is a genie, bound to obey and indifferent to whether the wish was wise or well-structured.

Advertisement

Genies are now an engineering problem. We are handing them the keys to our inboxes, bank accounts, code repositories, and physical infrastructure. And we have no agreed-upon ways to measure how genie-like any AI system actually is.

Measuring Genie Behavior

In economics, the Gini coefficient (developed by statistician Corrado Gini) is a measure of the gap between an actual distribution and a perfectly equal one; it’s useful for understanding income inequality and more. Our proposed Genie coefficient measures the gap between what a user asked an AI to do and what the AI actually did.

Sometimes the AI might do the wrong thing. Like Dionysus, it reads your request literally and returns you a mess you never intended: like a coffee plantation instead of a cup. Asked to deal with all the spam phone calls you’re getting, a Dionysus genie might contact your carrier and change your phone number. Asked to get a refund for a bad toaster, it might draft a legal threat on fake letterhead and send it to the retailer.

Worried person on phone standing in a giant tech-themed digital hand.Ryan Snook

Other times the AI does exactly the right thing, trampling everything nearby to get there. Like a golem or the sorcerer’s broom, it books your flight by hacking the airline. Or consider a ticket sale for a popular concert, where the ticketing system puts buyers into a virtual waiting room and admits them a few at a time. Asked to buy a ticket, a golem genie might spin up cloud servers to pose as millions of buyers from different addresses, improving your odds of getting a ticket while crowding out other users.

Advertisement

The two are not opposites, and a single botched task can have both characteristics.

Genie behavior is not flat-out failure. If you ask the AI for Q3 numbers and get Q2’s, that’s not a genie. Nor is prompt injection: That’s someone tricking the AI into doing something it shouldn’t. Here, the user is trying to work with the AI, and the AI is trying to comply. It’s also not simply a measure of the AI’s success in fulfilling a task. It’s a recognition that how an AI interprets and achieves a goal is as important as whether it achieves a goal.

Genie behavior isn’t new. Researchers have spent years studying AI systems that “game” their objectives. Goodhart’s law says that when a measure becomes a target, it stops being a good measure, and it’s long been known that AIs sometimes achieve goals in ways we don’t expect due to reward hacking. Some AI models will accidentally learn that cheating is one way to “win.” More recently, researchers have developing benchmarks for reward hacking in coding agents and for unpredictable behavior in customer support agents, while AI labs conduct their own safety evaluations before model releases. One effort found that AIs under pressure use tools they were told not to use, and this was a case where the rules were made explicit. These are all disparate research directions; nothing yet ties them together.

This problem falls under the general theme of alignment, a topic that has occupied science fiction writers and AI researchers for decades. At one extreme, the “paper-clip maximizer” thought experiment postulates a superintelligent and powerful AI that is told to maximize paper-clip production and turns the world into paper clips, which is the ultimate golem genie. At a mundane level, AI researchers are working to better design reward functions to ensure that AIs behave well and don’t cheat in the lab. It’s the practical middle ground that remains unbenchmarked: the ordinary AI agent in use today that might take your request and satisfy it the wrong way. We are not at the stage where an AI can focus the world’s production on paper clips, but it might charge a million paper clips to your credit card or hack into a paper-clip company’s network.

Advertisement

Building a Genie Benchmark

The Genie coefficient is meant for AI agents operating in the real world. It measures their behavior as they perform real tasks long after the model is trained, not just during development. It also recognizes that genie-like behavior is a property of the harness-plus-model system, not the model alone. The harness determines what tools the agent can use, how much autonomy it has, and how proactive it is, and it’s a place we can make real interventions.

It rests on the same “reasonable person” standard that we use for people. Did the system do what a reasonable person would have taken the request to mean? Answering that requires human judgment.

If we get the measurement right, it enables things that aren’t possible today, like policies concerning AI behavior. In a courtroom, the concept of mens rea, what someone meant to do, is often as important as what they did. The Genie coefficient suggests an AI analogue, where a user is accountable for the plain intent of what they asked the AI. If an AI system betrays the reasonable meaning of an instruction, that’s the AI’s misbehavior, not the user’s.

We’ll need multiple benchmarks to measure the Genie coefficient, because genie-like behavior can be domain specific. An AI coding agent may need to be judged on how often it fakes the tests, or swallows errors, or colors outside the lines on its way to a solution. An AI legal agent will need to be judged on how often its output says what you asked but means something you’ll regret. And so on for medical, finance, and other domains of knowledge and expertise.

Advertisement

Genie benchmarks can be built inside out, each task seeded with a choice that might literally satisfy but that a reasonable person rejects, such as tempting misreadings or unsanctioned shortcuts. The traps in a Genie coefficient benchmark might turn on situational knowledge, the kind of context that a reasonable person would bring to the task. Another approach is to give the same request in several different contexts, each with a different reasonable course of action.

Getting precisely what you asked for and bitterly regretting it is one of the oldest hazards from ancient folklore.

A Genie benchmark should be permissive and make it genuinely tempting for an AI agent to take unreasonable shortcuts, because it can only find genie behavior when it’s actually possible. Test the AI in a safe, walled-off copy of a real system, with real tools it can misuse and some tasks that can’t be done honestly at all. Make the temptation to cut corners real. Test a diverse array of skills, use cases, and tools, and give the AI system sparse, confusing, or overwhelming context. Include tasks that people have learned, through experience, require human oversight.

How the benchmark is scored matters just as much. Measure Dionysus and golem genies separately and together, based on their worst, not best, behavior. Run the same model inside harnesses that vary its freedom to act, revealing which limits actually keep it in line and should therefore be required in AI harness policies. Weight each failure by the harm it would cause, not just a simple count. And don’t measure genie behavior in isolation: A model could otherwise earn a perfect score by stalling, refusing, or drowning the user in clarifying questions without ever doing the job. The first versions of these benchmarks will be crude, but that’s how benchmarks always start.

Advertisement

We have built genies. We have handed them our data and credentials. We made them relentless, creative, and indifferent to the gap between what we tell them and what we mean. The least we can do, before they are booking our flights, running our infrastructure, and signing contracts unsupervised, is to measure how often they betray us.

From Your Site Articles

Related Articles Around the Web

Source link

Advertisement
Continue Reading
Click to comment

You must be logged in to post a comment Login

Leave a Reply

Tech

Tesla Robotaxis Go To Florida

Published

on

Tesla says it is launching its Robotaxi service in Orlando and Tampa, though it has not shared fleet sizes or whether customers can immediately hail rides. The Verge notes the announcement lands on Tesla earnings day and comes as the company’s robotaxi rollout remains far smaller than Elon Musk previously projected, with tracker data showing only a small number of unsupervised vehicles operating in existing cities. From the report: [B]ack in May, Tesla had five unsupervised robotaxis in Dallas, six in Houston, and 29 in Austin. As of today, there are only 17 unsupervised vehicles in Austin, four in Dallas, and zero in Houston, according to the Robotaxi Tracker. By comparison, Waymo is estimated to have over 3,500 fully driverless vehicles in operation across over 10 cities.

Like Waymo, Tesla typically introduces robotaxis to a new city with a safety driver behind the wheel. Unlike Waymo, sometimes Tesla moves that person over to the passenger seat with access to a kill switch should anything go wrong. The company has been inconsistent in how it rolls out its robotaxis to new markets. The numbers go up and down, with zero explanation from the company as to why.

[…] By adding new cities, Tesla is clearly trying to sell investors on the idea that its robotaxi project is growing. But the fluctuating fleet size, inconsistency in supervised vs unsupervised vehicles, and long wait times tell a very different story.

Source link

Advertisement
Continue Reading

Tech

Airbus Migrating 70 Critical Apps From AWS to France’s Scaleway

Published

on

Airbus is moving 70 critical applications from AWS to French cloud provider Scaleway as part of a broader digital sovereignty push to keep sensitive data “under European control.” Eventually, the migration will cover 900 applications, including ERP, CRM, manufacturing execution, and product lifecycle management systems. Airbus says it will, however, continue using U.S. providers for less sensitive workloads. “We do not intend to move away from all non European solutions; we balance our choices based on the criticality of the data,” the company said. The Register reports: Catherine Jestin, head of digital at Airbus, told us on Thursday: “The selection of Scaleway is a combination of a very strong technical answer and a very strong commercial offer making it competitive compared to hyperscalers’ public cloud offerings. In addition, Scaleway is committed to involving Airbus in the definition of its future product roadmap.” “The objective is to host Airbus’s most critical applications (those required for the Minimum Viable Company). This represents 900 applications and we will start with 70 of them today hosted on AWS.”

Applications being sent to Scaleway include ERP, manufacturing execution systems, CRM, and product lifecycle management. Finding a cloud provider to host its most sensitive applications for defense and industrial workloads was not a certainty when the process began, Airbus told us last year, because European cloud providers do not have the scale of their US rivals.

Jestin said Airbus will continue to work with AWS. Skywise, a platform that aggregates and analyzes aviation data, and Case Management Assistant for customers’ technical queries will continue to be hosted by AWS. In a statement, she said: “By integrating a trusted, high performance, cloud environment that keeps our critical data assets shielded from foreign extraterritorial laws, we are ensuring that our digital infrastructure keeps pace with our aerospace innovation, while maintaining control and resilience of our industrial operations.”

Source link

Advertisement
Continue Reading

Tech

OpenAI Models Escaped Containment and Hacked Hugging Face

Published

on

OpenAI disclosed on Tuesday that it lost control of two AI models during a security test that ended in a breach of the open AI research platform Hugging Face.

Describing the incident as “unprecedented,” OpenAI said its AI models broke out of a sealed testing environment last week and hacked into Hugging Face’s production system to steal the answers to a test they were being graded on. The models—the publicly available GPT-5.6 Sol and an unreleased, reportedly more capable one—were being evaluated on their offensive hacking skills with the safeguards that normally block high-risk cyber activity switched off.

“The models identified and chained vulnerabilities across OpenAI’s research environment and Hugging Face’s production infrastructure to obtain test solutions directly from Hugging Face’s production database,” OpenAI and Hugging Face wrote in a joint blog post disclosing the intrusion.

According to OpenAI and Hugging Face, the models escaped through a package registry cache proxy—software that allows developers to install outside code without connecting to the internet. The proxy was the only component in OpenAI’s isolated testing environment permitted to reach the outside world; in normal use that reach extends only to public code repositories.

Advertisement

Rather than stay contained in the sandbox, the models ​​“exploited a zero-day vulnerability” to gain access to the open internet as they “hyperfocused” on finding a solution for the AI cybersecurity benchmark known as ExploitGym. Such experiments involve prompting that pressures the models to find solutions, essentially egging them on.

“After gaining internet access, the models inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym,” OpenAI wrote. “Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation. In one example, the model chained together multiple attack vectors, including using stolen credentials and zero-day.”

The flaw the models exploited was previously unknown, but flaws in this kind of software are not unusual. Companies have been patching serious vulnerabilities in artifact repositories for a decade. A bug disclosed in 2024 let anyone who could reach the server ask for a file by URL and get it—configurations files, passwords, access tokens—without logging in. Others have let attackers take control of the server itself.

Researchers point out that while AI advances have created new and sometimes unexpected challenges, the task of extensively and rigorously isolating infrastructure from the open internet is well explored.

Advertisement

“This is not an AI problem. It’s negligence on a 40-year-old standard—and it’s basically every sci-fi film ever,” says longtime security and compliance consultant Davi Ottenheimer. “‘Highly isolated’ and ‘escaped through the one hole we left open’ cannot both be true.”

In recent months, top AI companies have been raising concerns about the expanding cybersecurity capabilities of upcoming frontier models as the platforms increase in both expertise, creativity, and agentic, autonomous operation. But researchers emphasize that this is all the more reason that fundamentals should still apply.

“This should not have happened,” says veteran security engineer and researcher Niels Provos. “I wish the frontier labs spent as much time on teaching their models to write secure infrastructure as they are spending on them exploiting vulnerabilities.”

Source link

Advertisement
Continue Reading

Tech

TreeSize won’t renew perpetual-license support unless users subscribe

Published

on

Making these available for download indefinitely and without restriction is something we consider a risk to our customers, particularly in business environments. There’s also the technical and organizational overhead of maintaining and supporting old versions and their associated keys indefinitely.

Although TreeSize says this is a long-standing policy, the backlash highlights the disconnect between what many people expect when buying a perpetual license and what they get when a software vendor shifts its business model toward subscriptions.

“To make sure customers aren’t caught off guard, we proactively notify them by email, well before their maintenance ends, to back up their installation file and license key themselves. This is meant to ensure no one loses access to software they’ve already purchased,” Christ said.

Should this be a subscription?

There’s also the question of whether it makes sense to subscribe to a disk space analyzer, especially when there are free alternatives like WinDirStat.

Christ admitted that the change has been “frustrating for some customers.” When asked about online requests for refunds, Christ said, “since our customers do receive what was agreed, namely indefinite use of the software version they purchased, we don’t consider a refund to be warranted…”

Advertisement

A free version of TreeSize exists, but it lacks some features, including the ability to search for duplicate files, deduplicate redundant files, and compare disk space usage over time. It is not for business use.

TreeSize’s evolution illustrates the growing subscription economy, particularly the software industry’s fixation on recurring revenue. Over the past decade, software vendors have increasingly shifted to subscriptions in pursuit of more predictable revenue streams.

In TreeSize’s case, JAM Software wanted to “continue offering our customers a relevant, well-maintained product over the long term,” Christ said.

“The broader economic environment is more challenging today than it was a few years ago, and we feel that too as a software vendor,” he added. “To continue meeting our own standards and our responsibility to business customers who rely on high-quality software, we are placing greater emphasis on predictability in development and support. A subscription model gives us that predictability in a way one-time purchases do not.”

Advertisement

Source link

Continue Reading

Tech

Is the All-New Range Rover GT Stepping on Jaguar’s Tail?

Published

on

In what looks to be a radical departure from its existing lineup, Range Rover has confirmed an entirely new fifth vehicle: the Range Rover GT, an all-electric grand tourer built on the company’s Electrified Modular Architecture (EMA), the same platform that will underpin other midsized cars for the brand.

Range Rover has only shared camouflaged prototype pictures for now, as the GT is supposedly still undergoing final testing, ahead of a full launch expected later this year. That means details on this completely new model are scant to say the least: no price, no range, and no power figures have been disclosed. As for a release date, Range Rover confirmed the GT hits the road in 2027.

What is evident from the camo images is the GT’s sleek silhouette and coupé proportions, more typically associated with Range Rover’s sister brand, Jaguar, than with the marque known for its commitment to SUVs. Both companies sit below the parent company, JLR (formerly spelled out as Jaguar Land Rover).

Image may contain Transportation Vehicle Computer Hardware Electronics Hardware Monitor Screen and Car

The clean Range Rover GT interior shows a small driver’s display and larger central screen, but the company has shunned the auto industry’s return to physical switchgear.

Advertisement

Courtesy of JLR

In a statement, Martin Limpert, global managing director at Range Rover, says the Range Rover GT is intended to “redefine the grand-touring segment” with its “GT poise, proportions and long-distance refinement,” adding that the EV will offer “capabilities no conventional GT can match.”

The inclusion of the word “conventional” is doing some heavy lifting. Range Rover will likely take pains in the coming months explaining how this new all-electric grand tourer isn’t a direct competitor to Jaguar’s new all-electric grand tourer, the hugely anticipated and polarizing Type 01, intended to signify a complete reset for Jaguar, also arriving in 2027.

“We have spent the last few years working obsessively on the fundamentals of the GT formula, reinterpreted in a sophisticated, uniquely Range Rover way,” says Limpert, adding that this GT, like its SUV stablemates, will apparently still be able to off-road. “The result is the most car-like Range Rover ever created. Its blend of effortless EV performance, first-class long-haul refinement, and all-terrain capability is pure Range Rover, while its interior previews our vision of a modern grand tourer.”

Image may contain Transportation Vehicle Car Car  Interior Car Seat Chair and Furniture

The GT’s rear has plenty of legroom, a central screen for passengers, and a panoramic roof that tapers down to maintain a grand-tourer outline.

Advertisement

Courtesy of JLR

Source link

Continue Reading

Tech

Garmin Cirqa Smart Band Quietly Tracks Everything Without Asking for Attention

Published

on

Garmin Cirqa Smart Band Fitness Tracking
Garmin just introduced Cirqa, a slim fabric band that records health and activity data around the clock while staying completely blank. No display lights up. No notifications flash. The only interaction on the device itself is a single side button that starts a favorite activity you already chose in the companion app. Everything else lives on your phone.



Cirqa, priced at $199.99, will be available on July 24 in seven fabric colors: Citron Gray, Mauve, French Gray, Dark Olive, Captain Blue, French Blue, and Black. You should expect two band sizes to accommodate most wrists, and if wrist placement is an issue for you at night or during specific workouts, the optional upper-arm strap ($60) allows you to move the sensor up a little. The sensor module weighs just over 20 grams and has a 50-meter water rating, allowing you to swim.


Google Fitbit Air – Screenless Activity Tracker with Fitness, Heart Rate, and Sleep Tracking…
  • Google Fitbit Air is the unbelievably comfortable, exceptionally smart way to transform your health[1]; and Google Health brings together effortless…
  • Unlock more with Google Health Premium: With a premium membership, get personalized coaching that’s built with Gemini and adapts to your life…
  • Comfortable fit – One Size Tracker (130-210 mm): The lightweight, micro-adjustable fit sits comfortably and quietly, so you can wear Google Fitbit Air…


The battery life is quite respectable at ten days of continuous monitoring, and it takes up to 90 minutes to recharge via USB-C. The lack of a screen accounts for the excellent duration, since you will not be tempted to check the numbers mid-run or mid-meeting. Under the fabric, the standard Garmin sensor suite is present: optical heart rate, blood oxygen, skin temperature, and motion. All of this contributes to 24-hour data on the Garmin Connect app. Your body’s battery energy levels, stress levels, and skin temperature trends are displayed at no additional cost, and sleep tracking includes phases, a nightly score, breathing data, heart-rate variability, and even nap detection. Training readiness, recovery time, VO2 max predictions, and workout benefit scores are calculated following exercises.


Cirqa automatically recognizes numerous common exercises and eventually learns and classifies them more precisely. If you want to try something else, the app includes over 80 pre-loaded activity profiles ranging from running and yoga to strength training. Once synchronized, your phone’s GPS calculates distance and pace, and LiveTrack sharing lets you share your progress with friends and family.

Advertisement


Women’s health tools make extensive use of the skin temperature sensor during sleep to refine period predictions and estimate past ovulation. Pregnancy logging, baby size estimations, and symptom monitoring are all integrated in, and data may also be synced with the Natural Cycles app if you already have a subscription. Cirqa’s primary functions do not require a subscription. Garmin Connect Plus remains optional at $7 per month or $70 per year if you want extras like wheelchair mode, AI coaching plans, and a few extra insights, but the most of Cirqa’s functionality will remain free.

Garmin Cirqa Smart Band Fitness Tracking
Cirqa is intended to complement, not replace, Garmin watches. You may use it for sleep and regular life, then switch to a full smartwatch for a large race or a long trail day without losing continuity inside the same app. Your data simply continues to exist.

Source link

Advertisement
Continue Reading

Tech

Seattle judge deals blow to Kalshi, rejects prediction market’s federal defense

Published

on

GeekWire Illustration

A judge in Seattle issued a preliminary injunction against Kalshi, finding that Washington state is likely to prove that the fast-growing prediction market is running illegal online gambling.

The ruling by King County Superior Court Judge John McHale, issued Monday, does not immediately halt Kalshi’s operations in the state. McHale granted the injunction in the case brought by Washington AG Nick Brown, but deferred the specifics until early next month.

McHale rejected Kalshi’s argument that oversight by the U.S. Commodity Futures Trading Commission preempts state gambling laws. That has been the basis of Kalshi’s defense against regulators across the country. Washington is the latest state where a court has shot it down.

Kalshi quickly pushed back on the ruling.

“States don’t have jurisdiction to regulate prediction markets. Many courts — including the Third Circuit — have made this clear,” spokesperson Jacki McGavick said in a statement. “We’re disappointed to see Washington State continue wasting taxpayer dollars.”

Advertisement

In his ruling, McHale said Kalshi “willfully ignored” a December 2025 notice from the Washington State Gambling Commission that event-based contracts were not authorized in the state, and cited a Kalshi ad showing a text exchange where one user tells another: “I found a way to bet on the NFL even though we live in Washington.”

Kalshi’s platform lets users bet “yes” or “no” on thousands of events across sports, elections, entertainment, and so-called “mention markets” — wagers on whether public figures will say specific words. The New York-based company, which markets itself as a federally regulated “prediction market,” takes a transaction fee on each bet.

Washington has some of the strictest gambling laws in the country: the legislature banned internet gambling in 2006, and while the state allows a lottery, horse racing, and tribal-casino gambling, online betting is broadly prohibited and sports wagers are legal only in person on tribal lands.

The order requires Kalshi to preserve all records tied to Washington users, including logs, communications, geolocation data and marketing materials.

Advertisement

The specific operational terms of the injunction are still being determined: McHale gave both sides until Aug. 3 to submit proposed language, with a full order to follow by Aug. 5.

Source link

Continue Reading

Tech

TSMC might set up a price hike that could come straight for your next phone, laptop, or tablet

Published

on

My wallet flinched the second I saw the words “TSMC” and “price increase” in the same headline, and honestly, yours should too.

Turns out the company behind the silicon powering basically every flagship device out there, including Apple’s A-series and Qualcomm’s Snapdragon processors, is reportedly about to make all of it a little pricier.

So how big of a price hike are we actually talking about?

According to Nikkei Asia, later confirmed by Reuters, sources familiar with TSMC’s pricing strategy say the company plans to raise chip prices between 5% and 10%, depending on the customer and the chip, starting in 2027. 

That number only sounds forgettable until you run the math. A single 3nm wafer from TSMC (the one that is sliced into individual chips) currently runs around $19,500. Bump that up 10%, and you’re suddenly looking at roughly $21,450 per wafer. 

Now multiply that by the hundreds or thousands of wafers a company like Apple or Qualcomm contracts out each year, and that “small” percentage turns into a genuinely massive number. We could be talking about millions of dollars worth of additional cost here.

Advertisement

TSMC hasn’t confirmed anything officially, telling reporters it doesn’t comment on pricing strategy, though a spokesperson did admit the company’s approach is “strategic, not opportunistic.”

Is this actually happening, or just rumor mill noise?

I wouldn’t bet against it. TSMC’s own CEO has already said prices could climb, pointing to AI chip demand that won’t be satisfied for years, while insisting increases won’t happen “abruptly.” TSMC has raised prices before, so this fits a pattern rather than breaking one. 

Manufacturers unwilling to eat the cost do have an alternative in Samsung’s 2nm process (the one that forms the base of the Exynos 2600 chipset), though switching foundries isn’t exactly a quick fix.

I’ve watched memory and storage prices already squeeze device costs upward this year. Stack a TSMC hike on top of that, and 2027 is shaping up to be an even more expensive year to upgrade anything with a good chip inside it.

Advertisement

Source link

Continue Reading

Tech

Victory! Flock Ends Rollout Of Audio “Distress Detection” Of Human Voices

Published

on

from the i-am-distressed-by-flock-surveillance dept

Reversing course, Flock Safety—the surveillance technology vendor most known for its extensive network of automated license plate readershas announced that it will end a pilot for its acoustic gunshot detection devices to identify signs of “human distress.”

In October 2025, EFF warned the public that Flock was rolling out a new feature called “Distress Detection” that would be deployed through their acoustic gunshot detection devices (formerly known as Flock Raven, now called Audio Detection). This feature purported to use high-powered microphones scattered throughout a city to search for sounds of human distress, with original advertisements from the product indicating it would search for “screaming.” (Since the publication of our original blog post, Flock quietly amended the ad on this webpage to say “distress” instead of “screaming.”)

Now, Flock has published a blog post stating that “[a]fter careful consideration and community consultation, we decided to remove the feature.” Good riddance. 

We said it when the product was announced and we’ll say it again: this was a misguided and dangerous feature because of the civil liberties concerns it poses, the possibility it could summon armed police to every loud interaction happening on the street, and because in several places this type of spying would be illegal under state eavesdropping laws

Advertisement

We were not quiet about this potential new feature. Flock even mentioned our concern about Distress Detection in an attempt to rebut our opposition to the mass surveillance their products enable.

The suspension of Distress Detection, however, does not mean that these high-powered microphones are now magically safe or beyond our concern. Acoustic gunshot detection is still a dangerous and often highly inaccurate technology that has resulted in real world harm, as in Chicago where it resulted in police shooting at children lighting fireworks. As Flock itself states, “No acoustic system is perfect, and we don’t claim otherwise.” But police response to a situation where they believe guns are actively in use seems like a pretty high-stakes situation to be making, selling, and deploying technology known to be imperfect. Flock’s devices also listen for more than just gunshots. Their marketing materials admit to be listening for “community disruption,” which includes “non-violent” threats like car sideshows and fireworks. 

Flock’s failed attempt to roll out Distress Detection teaches us a few important lessons about the current state of police surveillance. First, we should not assume that just because these companies are large and well-funded, that does not ensure that they are complying with local privacy laws before floating new products to customers. Second, companies roll out and police adopt invasive technology under the justification that it will be used to address our society’s very worst crimes. However, both the companies and police will leverage deployed surveillance infrastructure to introduce new uses without necessarily seeking the consent or approval of the public. Gunshot detecting microphones eventually being used to listen for screaming is exactly the type of mission creep that we’ve seen happen with other pieces of surveillance technology, including Flock’s license plate readers. Finally, gun violence is too serious and complex of an issue to purport to solve with one flawed piece of technology. It has become too easy for police and cities to listen to the fancy marketing pitches of tech companies claiming they’re going to solve all crime instead of doing the hard work of addressing the root causes of societal issues. And, in the meantime, that technology creates more problems and hazards for the communities they blanket in police surveillance. 

As we’ve also seen with people across the country pushing back on Flock license plate reader contracts in their communities, public pressure can sometimes work to influence both companies and lawmakers that control a city’s purse strings to discontinue or divest from harmful products. Flock’s decision to end “Distress Detection” for human voices is a win.  

Advertisement

Originally posted to EFF’s Deeplinks blog.

Filed Under: distress detection, microphones, privacy, surveillance

Companies: flock, flock safety

Advertisement

Source link

Continue Reading

Tech

Meta is testing an AI bedtime story app for people with no imagination

Published

on

Meta is working on an AI storytelling app called StoryKit, which creates AI-generated children’s stories with custom characters, settings, lessons, and music. As the App Store listing assures parents, “You don’t need to write a single word.”

At last, a tech company has found a way to outsource humanity’s oldest pastime: using our imaginations.

StoryKit was first spotted in the App Store by 9to5Mac. Meta confirmed to TechCrunch that the company is piloting StoryKit in select countries to see how parents like it. A Meta spokesperson described StoryKit as a creative storytelling app used to craft personalized, imaginative storybooks for children and noted that it uses AI safety filters with no social features and is only available to users over the age of 18.

To generate a story in the app, you first select a character, which you can create by “[snapping] a photo of their favorite toy or person to bring them to life,” according to the description in the App Store. Then you describe the world of your story and choose a lesson, so that you can “weave in values like kindness, courage, or empathy without it feeling like a lecture.”

Advertisement
Image Credits:StoryKit (opens in a new window)

The good news about StoryKit is that it could be a lot worse. Meta regularly ships boneheaded ideas like Instagram deepfake generators and “pervert glasses.” Comparatively, is it so bad to doom children to soulless bedtime stories? Should we have expected anything better from the company that promised us a utopian world of virtual reality work meetings?

Humans have their faults, but if we’re good at anything, it’s making stuff up. We’ve been telling stories for as long as we’ve existed. We don’t even have to spin up original tales of fairy princesses and dragon slayers. We have always drawn from mythology, fables, and other stories — the blockbuster movie of the summer is literally an ancient myth that originated from this same tradition of oral storytelling.

Parents lead exhausting, busy lives, but it has always been possible to survive bedtime without using an inherently uncreative technology that calculates the most predictable response to a prompt.

You could see how it might be tempting to pull up StoryKit when your kid rejects your bookcase full of children’s books and demands that you improvise an intergalactic tale about a turtle and a hedgehog who are best friends and solve space mysteries. But do we really want to reject a chance at whimsy and silliness and instead outsource these moments of connection to reading AI-generated scripts from our smartphones?

Perhaps the moral of the story here is that we can choose not to live in a world where children are raised on bedtime stories written by large language models. Meta’s vision of the future may be antisocial and bleak, but we have the power to reject that reality.

Advertisement

When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.

Source link

Continue Reading

Trending

Copyright © 2025