Connect with us

Tech

Reddit aims to make ‘karma’ less important for first-time posters with shift to AI moderation tools

Published

on

Reddit on Wednesday announced a series of changes to its infrastructure and tools, designed to make it easier for people to participate on its site, protect against scraping and spam, and aid in the moderation of its online communities. In addition to providing a new suite of moderation tools, the company said it’s working on more advanced abuse prevention systems that could eventually help communities move away from using things like account age and “karma” to determine who’s allowed to post.

Karma, Reddit’s digital reputation system, allows users to raise their score by posting and commenting helpful, friendly, or funny responses that lead to upvotes. Originally designed to weed out spam, bots, and trolls, many communities came to rely on karma and other account-age restrictions that make it difficult for legitimate newcomers to participate in their communities. Reddit says it now wants to shift to stronger, built-in abuse prevention systems to do more of the moderation work, so communities can be more open to new users.

Reddit didn’t say karma would go away entirely with these coming changes, but it did suggest that its importance could dwindle in the future.

“This will make it easier for genuine new users to participate and easier for mods to welcome them with confidence,” Reddit’s announcement stated.

Advertisement

Related to this, the company said it’s expanding the test of its new suite of AI-powered moderation tools, the Rules Hub, which helps moderators by choosing when to automatically enforce certain community rules and what action should be taken when doing so. Initially used by 700-plus communities, Reddit says that all new communities can now test the tool. Later this year, the tool will be widely available to all communities, both new and existing.

With Rules Hub, Reddit uses large language models (LLMs) to determine whether a post or comment matches the intent of a rule, which the company says allows it to “better handle nuance,” natural language, and edge cases.

Reddit also expanded the capabilities of other tools, including one that helps moderators inform users about their community’s rules when posting and commenting, and another that helps members set their flair — a custom tag users append to their name that only appears in the community they’re posting in.

“Today, new users can encounter invisible barriers like account age and karma thresholds, unclear removals, and poor community discovery,” Reddit explained. “We want new users around the world to be able to easily find relevant communities, understand their rules and norms, and make useful contributions.”

Advertisement

The company also said it would make changes to its legacy desktop site, Old Reddit, by shifting its moderator workflows to its new stack and migrating other critical bots.

On Reddit’s second-quarter earnings calls with investors last week, company executives stressed that one of Reddit’s near-term goals was to convert its half-billion weekly active users to daily users through product updates. The company had crushed earnings with revenue of $805 million and earnings per share of $1.25, above estimates, but still saw its stock sink because of what Reddit CEO Steve Huffman described as “choppy” search referral traffic.

When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.

Source link

Advertisement
Continue Reading
Click to comment

You must be logged in to post a comment Login

Leave a Reply

Tech

The Most Dangerous AI Hacking Techniques Still Have Humans in the Loop

Published

on

Agentic AI has permanently changed cybersecurity by making it quicker and easier to discover vulnerabilities in software and fix them—or develop so-called exploits to weaponize them. But longtime web security researcher James Kettle wanted to look beyond the bug-hunting apocalypse to explore a question that has taken on even more urgency as major AI organizations disclose real-world examples of rogue AI hacking: Can agentic AI develop novel, abstract hacking methods, from concept through to practical attacks?

At the Black Hat security conference in Las Vegas on Wednesday, Kettle presented his findings, which illustrate both AI’s rapidly advancing cybersecurity capabilities and its limitations. For now, the answer to Kettle’s question is nuanced. He concluded that AI is perhaps minimally capable but extremely limited in its ability to devise new attack paths in a fully autonomous way. Importantly, though, when paired with human guidance and insight in key moments, Kettle found that AI is an extremely powerful partner in conceptualizing and uncovering new strategies for hacking.

After spending years researching web security vulnerabilities, Kettle now says he has uncovered an entirely new area of potential vulnerability—dubbed Shared-Parser Confusion—as the result of an AI revelation about web servers using shared code to process both requests and responses.

“This is an absolutely massive deal, because if you think about it, requests to a website are completely untrusted, they could be anything, but responses are trusted,” Kettle told WIRED ahead of his conference talk. “So this is a major attack surface and potentially spills into a lot of different attack types.”

Advertisement

The finding came out of months of experiments that began in September 2025 using Anthropic and OpenAI’s latest models at the time. Kettle wanted to explore AI’s ability to do theoretical security research, but quickly realized that one obstacle was that the systems were attempting to pass existing research off as original by returning findings about extremely esoteric topics that were difficult to vet. With this in mind, he decided to scope his tests more narrowly so the AI systems were working within his own area of web security expertise. This way he had total command of the material and knew that AI couldn’t trick him. Additionally, Kettle realized that by synthesizing his own research methodology and training models on it, he could probe deeper into what the systems were capable of extrapolating on their own.

“I’m interested in pushing AI to the absolute limit to see where it fails and where you need a human,” Kettle says. “There are still very few people talking about where the limits are, especially in the security space, because there aren’t incentives to talk about that angle. Everyone wants to be seen as AI native, not talk about where their system falls apart completely.”

As Kettle honed his experiments—providing models with more methodological data and more refined parameters—and as time passed and more powerful models debuted, he says the systems had more and more findings at a rate far surpassing his own, creating what he describes as a productive research feedback loop.

“It was really interesting going through the process. It would have notable findings maybe every two days without me even logging into the system, to the point that it was making me anxious,” Kettle says, “like I almost don’t want to know. It was so many research leads that you have FOMO about not exploring all of them, so it forces you to automate more analysis.”

Advertisement

In addition to finding more proven examples of certain vulnerabilities in a few months than Kettle could likely find in a few years, he also hoped that the AI system could find an entire novel class of those types of bugs. And in a way it did succeed, he says, but the finding related to an extremely rare type of bug and was not actually exploitable in the one vulnerable target available. Kettle emphasizes, though, that the Shared-Parser Confusion finding was so significant, even though it was a human/AI collaboration, because it illustrates the reality of how AI systems can contribute most powerfully to cybersecurity work right now for both defensive and offensive hacking.

“It wasn’t able to prove this itself, but it analyzed some real, proven findings and came up with the hypothesis, and I evaluated it and confirmed it,” Kettle says. “That’s probably going to be the discovery that has the biggest long-term impact. It couldn’t do that on its own, but I would never have found that on my own for sure. Even if you gave me the single line from the [documentation], I wouldn’t have seen it. But together we managed to find it.”

Source link

Advertisement
Continue Reading

Tech

Jeff Dean and other top AI researchers are leaving Google to launch their own startup

Published

on

Jeff Dean, one of the longest-serving and most influential executives at Google, is stepping down from the search giant to launch his own AI startup.

Coming with him as co-founders are several other top researchers at the company, including Sanjay Ghemawat, a top engineer and senior fellow at Google; Quoc Le, a key AI researcher and founding member of Google Brain; and Oriol Vinyals, a senior research scientist at Google DeepMind. Dean reportedly plans to serve as CEO.

Together, the group is starting Discovery Loop, a public benefit corporation that seeks to use AI to turbo-charge scientific research.

Discovery Loop says it plans to use high-octane algorithms to initiate and iterate thousands of experiments simultaneously, with the goal of partially automating the research process and expanding the scale at which experimentation can be conducted.

Advertisement

The startup is also interested in using AI to help create more powerful AI (a process known as recursive self-improvement), which would cut human iteration out of the loop entirely.

“While science and engineering have tremendously advanced society over past centuries, progress has traditionally relied on slow, sequential human iterations, creating a significant bottleneck,” the company said in a press release. “Discovery Loop is developing advanced AI systems that leverage massive computational scale to fundamentally transform the speed and efficiency of innovation by automating complete experimental loops.”

Using AI to accelerate scientific discovery has been a key interest of the science and tech communities for many years, but until recently, it remained a largely experimental field with limited commercial application.

The company has received financial support from a number of different sources, including Google’s parent company Alphabet. The initial funding round is being co-led by Radical Ventures and Khosla Ventures, the company announced Wednesday. Kleiner Perkins, Lightspeed, and Doerr Capital also participated.

Advertisement

“The next great frontier for AI is to go beyond answering questions and to begin making discoveries,” the founding team said in a joint statement. “By fundamentally accelerating how engineering and scientific discovery are conducted, we can deliver the benefits of transformative technologies to the world far sooner.”

Dean has worked at Google since 1999 and was Google’s 30th employee. Over the years, he has contributed to Google search’s core infrastructure, including its crawling and indexing system and its query-serving system. He has also had a significant impact on Google Gemini’s multimodal models and played a leading role in the company’s early AI research.

“We think there is opportunity for AI to more fully automate what has traditionally been a very human-intensive experimental loop,” Dean told the New York Times. “You will get both a higher quantity and a higher quality of experiments, and that will lead to scientific breakthroughs and advances.”

When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.

Advertisement

Source link

Continue Reading

Tech

Get up to 37 days of battery life for less with this Garmin Fenix 7X Pro deal

Published

on

£501 off a watch built to survive years of proper use isn’t a discount you see very often, and it’s exactly what’s happened to the Garmin Fenix 7X Pro Sapphire Solar Edition right now.

The Fenix 7X Pro Sapphire Solar Edition has dropped from £870 to £369, a saving of 58% that turns Garmin’s most capable outdoor watch into something you can actually justify buying instead of quietly wishing for while telling yourself the current watch still works fine.

Deal Garmin fenix 7X PRO SOLARDeal Garmin fenix 7X PRO SOLAR

Get up to 37 days of battery life for less with this Garmin Fenix 7X Pro deal

Get up to a huge 37 days of battery life for a lot less with this rare discount on the Garmin Fenix 7X Pro.

Advertisement

View Deal

That kind of saving matters most if you’ve spent months eyeing a proper outdoor watch but couldn’t quite stomach paying full price for something you’d mostly use to check the time and count your steps around the house.

Advertisement

None of that applies once you’re actually out using it, and the fenix 7X Pro tracks heart rate, stress, sleep and Body Battery energy levels continuously, so you get a genuine read on how your body is coping rather than a rough guess.

If you’re curious how the 7X Pro compares, our tech expert Michael Sawh tested the standard Fenix 7 Pro and found much the same core experience, praising the LED flashlight, upgraded heart rate sensor and new Endurance and Hill Scores as solid rather than groundbreaking additions.

Advertisement

That same toughness carries over to the X, though it goes further outdoors with multi-band GNSS, topographic mapping and a titanium bezel built to survive the kind of terrain that would scratch up a normal smartwatch within a single afternoon hike.

Advertisement

The Whatsapp LogoThe Whatsapp Logo

Get Updates Straight to Your WhatsApp

Join Now

Advertisement

The number that stands out most is battery life, because the solar charging built into the display keeps the fenix 7X Pro running for up to 37 days in smartwatch mode or 122 hours with GPS switched on.

That’s the sort of number that changes how you actually think about wearing a watch, since charging it becomes something you plan around once a month rather than something you have to remember to do every single night.

The everyday side hasn’t been left out either, with smart notifications, Garmin Pay for contactless payments and safety tracking features that quietly stay switched on whether you’re deep in the mountains or just walking the dog around the block.

So if you’ve been putting off the upgrade because the full price never quite made sense to you, doesn’t a 58% cut on the Garmin fenix 7X Pro Sapphire Solar Edition finally answer that question for good?

Advertisement

Advertisement

SQUIRREL_PLAYLIST_10148964

Source link

Advertisement
Continue Reading

Tech

The Disney Plus App Is Getting TikToks. Is a Free Plan Next?

Published

on

Disney CEO Josh D’Amaro confirmed Wednesday that the entertainment giant is considering adding a free option for Disney Plus as a way to attract more people to the streaming service.

“We’re exploring a free product for consumers,” D’Amaro said during the company’s 2026 third-quarter earnings call. “A free offering could help us drive top-of-funnel Disney Plus subscriber growth.”

Also on Wednesday, Disney and TikTok announced that they have struck a deal to bring Disney-themed TikToks, made by selected creators, into the Disney Plus app. These videos will be in the app’s new Verts section and will be rolling out in a pilot program in the US later this year.

The House of Mouse’s hints about a potential free tier appear to confirm a report from Business Insider last month. The rumored new option would help Disney keep its ground in an ever-changing streaming landscape. In June, Fox announced that it is acquiring Roku for $22 billion

Advertisement

Meanwhile, the ongoing sale of Warner Bros. Discover to Paramount SkyDance (itself created from a merger only last year) has sparked protests by entertainment industry workers about consolidations, amid concerns about job losses, fewer projects produced and for newsrooms, less editorial freedom. David Ellison, CEO of Paramount SkyDance and son of tech tycoon Larry Ellison, has rejected these concerns.

Disney executives put a rosier spin on things.

“A more consolidated industry is really a better investment backdrop,” D’Amaro said, “and we have a long history of partnering and streaming, and we believe that we can just keep building on that.”

Disney has its own Plus app, Hulu and ESPN, and has deals with Paramount-owned HBO Max, fueling its 11% increase in streaming-related earnings in the third quarter to a total of $5.53 billion. But free live TV streamers, like Fubo, are enticing viewers who don’t want to break the bank to watch individual live events. That’s where a free tier for Disney Plus may make sense.

Advertisement

Disney’s world beyond streaming

“It’s become clear Disney Plus is no longer just a streaming service, and perhaps that’s a good thing,” said Mike Proulx, research director at Forrester. Streaming “has a customer relationship problem,” which Disney is trying to solve by making its streaming service the focal point of its digital offerings. 

“Movies and series are the foundation, but live sports, games, merchandise, creator content, and future perks give consumers more reasons to stay engaged with Disney,” Proulx said.

This is certainly what Disney is hoping for. D’Amaro defended recent box office flops Star Wars: The Mandalorian and Grogu and the live-action Moana. He said that while these didn’t meet box office expectations, the movies prompted growth in other parts of the business.

For example, The Mandalorian and Guru barely broke even, making it the lowest-grossing live-action Star Wars flick. But it also prompted a surge in merchandise sales – like apparel and Lego sets – and prompted visits to Walt Disney World’s Millennium Falcon attraction. The original Moana movie is one of the most-streamed movies on Disney Plus, and D’Amaro said Disney expects its live-action companion to perform well, despite reportedly standing to lose $100 million at the box office.

Advertisement

Then there’s a franchise like Toy Story, whose latest installment this summer was an undeniable theatrical success and also supported Disney’s other businesses. 

“The nature of the film industry is such that it is more of a portfolio game,” D’Amaro said. “The theatrical window, in a lot of ways, it’s just one data point. And the real value of that IP is the cumulative benefit of decades-long storytelling and our ability to take that IP and play it into the entirety of the Disney flywheel.”

Source link

Continue Reading

Tech

Minecraft Is Coming To Switch 2 On October 27

Published

on

There’s no confirmation of Joy-Con 2 mouse control support just yet.

The best-selling video game of all time is making its way to yet another platform, as Minecraft will land on Nintendo Switch 2 on October 27. Developer Mojang says there will be an upgrade path for those who own the Switch version. More details about that will be revealed later.

Minecraft on Switch 2 should have stronger visuals than the the previous-gen version. By default, it will enable Vibrant Visuals, an enhanced graphics mode that Mojang released last year. It features directional lighting, volumetric fog, enhanced shadows and more. The Super Mario Mash-Up pack that’s exclusive to Nintendo versions of Minecraft will receive a Vibrant Visuals boost as well. (Vibrant Visuals won’t be available in split-screen play, some other game modes and certain Marketplace Worlds and other content.)

Mojang hasn’t confirmed whether the Switch 2 version will support Joy-Con 2 mouse controls, which could help players create more precise builds faster. We’ll surely learn more about the port in the coming weeks.

Advertisement

The Switch 2 version of Minecraft will arrive soon after the latest spin-off in the series. Minecraft Dungeons II will hit PC and consoles (including Switch and Switch 2) on September 29.

Source link

Advertisement
Continue Reading

Tech

Google to cut 52 employees in Washington state

Published

on

Google’s Kirkland Urban campus in Kirkland, Wash. (GeekWire File Photo / Kurt Schlosser)

Google is laying off 52 employees in Washington state, according to a state filing released Wednesday.

The job cuts include software engineers, engineering and product managers, mechatronics engineers, UX designers and recruiters located in Kirkland, Redmond and Seattle offices as well as remote employees.

The layoffs are scheduled to take effect in September and early October. The notice from Google filed with the state’s Employment Security Department did not provide an explanation for the cuts.

GeekWire has reached out to Google for comment and will update if one is forthcoming.

Business Insider reported layoffs in the Google Cloud division earlier this summer, but did not say how many jobs were impacted.

Advertisement

While Microsoft, Amazon and other tech companies have cut Washington employees in multiple rounds over the past year, this is Google’s the first significant reduction in the state since 2023. In January of that year, Google parent Alphabet shed 12,000 workers globally or 6% of its workforce. It did not indicate how many jobs were lost in Washington.

Washington is one of Google’s largest engineering hubs outside of its Bay Area headquarters. It opened offices in the state more than 20 years ago, with a focus on Android, Chrome, Cloud, Maps, Ads and other projects.

Zillow Group this week disclosed it is eliminating 91 jobs in Washington state, landing heavily on senior staff.

Source link

Advertisement
Continue Reading

Tech

IBM’s agentic AI platform is under active attack

Published

on

security

A critical Langflow flaw allowing RCE on default deployments is being exploited, says the CISA

A critical vulnerability in IBM-owned, low-code AI builder Langflow lets unauthenticated attackers execute code remotely on vulnerable default deployments, potentially putting organizations running those instances at immediate risk.

The Cybersecurity and Infrastructure Security Agency (CISA) on Tuesday added CVE-2026-9198 to its Known Exploited Vulnerabilities catalog after identifying evidence of active exploitation and urged organizations to apply the vendor’s mitigation guidance as soon as possible. IBM says the flaw affects Langflow OSS versions 1.0.0 through 1.10.0 and recommends upgrading to version 1.10.1 or later; at the time of writing, the most recent version is 1.11.2.

Advertisement

Langflow, for those unfamiliar, is one of the more accessible AI agent builders on the market, as our hands-on look at the tool earlier this year demonstrated. It’s available on Linux, Windows, and macOS, and is basically an end-to-end, drag-and-drop GUI where users can construct agent workflows without having to know much, if anything, about the underlying code.

IBM owns the platform now, but Langflow was originally developed by Logspace, which was acquired by DataStax in 2024 before IBM scooped up DataStax, and Langflow with it, in 2025. 

The acquisition of DataStax and its tools like Langflow by IBM paved the way for Langflow to be integrated into watsonx.ai, IBM’s AI development studio, as a piece of middleware extending watsonx.ai’s capabilities.  

The ownership changes, however, didn’t stop the critical flaw from making it into production releases before it was finally fixed.

Advertisement

According to IBM, the vulnerability affects default Langflow deployments and combines two issues that, when chained, allow an unauthenticated attacker to execute code remotely.

First, there’s the matter of an auto-login endpoint in default deployments that’s willing to mint superuser tokens to any network caller. Combine those easily obtained superuser rights with the second issue, a code validation endpoint that’ll run any old Python code thrown at it, and you’ve got a recipe for someone taking over your entire Langflow server, or worse.

The CVE itself was published on July 17, meaning that it hasn’t taken long for bad actors to realize what they could do with RCE on any system hosting a default Langflow deployment with auto login enabled and that code validation endpoint left accessible on a network. 

Langflow itself isn’t a vibe-coding platform, instead serving as an interface for building agentic and RAG workflows, so don’t blame vibe coding or no-code security failures for this one. Instead, what we appear to have is a standard case of how default configuration deployments can easily be a disaster.

Advertisement

It’s unknown how extensively exploited this vulnerability is; we’ve reached out to IBM to learn more. ®

Source link

Continue Reading

Tech

Rare Soft X-Ray Flash Gives Scientists a Clear View of a Star’s First Explosive Moment

Published

on

Soft X-Ray Flash Supernova SN 2026gzf
Soft X-rays from a galaxy about 500 million light-years away triggered an urgent worldwide search in March 2026. The Einstein Probe satellite, built by the Chinese Academy of Sciences with the European Space Agency, recorded a brief pulse of these X-rays and labeled the event EP260321a. Ground telescopes responded within an hour and quickly found a supernova growing brighter by the minute. Researchers later named the explosion SN 2026gzf.



Two astronomy teams were diving into the same region, utilizing a range of NSF NOIRLab resources. One, led by Brendan O’Connor of Carnegie Mellon University, and the other by Jillian Rastinejad of the University of Maryland, reached the same conclusion after evaluating the data. Their early observations suggested that the explosion began with a shock breakout. A shock breakout occurs when an immensely powerful shockwave from a collapsing core smashes through the star’s surface, allowing the supernova to shine. This isn’t uncommon in supernovae, but they’re difficult to detect because they only endure a few seconds to a few hours. To put it in context, there has only been one validated x-ray shock in the last 20 years, and scientists were certain this was the real deal. EP260321a is a once-in-a-lifetime scientific discovery.


LEGO Icons NASA Artemis Space Launch System – DIY Rocket Model Building Set for Adults, Ages 18+ – Gifts…
  • NASA rocket model kit – Launch into a creative project with the LEGO Icons NASA Artemis Space Launch System model building project for adult space…
  • What’s in the box? – This creative building set includes everything you need to craft a multistage rocket with 2 solid-fuel boosters, an Orion…
  • Features and Functions – This NASA-themed rocket model features retractable launch tower umbilicals, rocket support and crew bridge, detachable…


People looking over the data discovered that SN 2026gzf is a broad-lined Type Ic supernova. These explosions are typically extremely intense, resulting in streams of material that travel at the speed of light. They are frequently accompanied by a gamma ray burst, which is one of the most tremendous energy outputs ever observed. Unfortunately, in this case, the teams were unable to identify a gamma ray burst, its high-speed jet, or even the afterglow. O’Connor hypothesized that the jet was likely muted by the star’s surface or surrounding material before breaking loose. Even though the explosion was quite intense, the resulting x-ray flare was exceptionally mild for something linked with a broad-lined Type Ic supernova.

Soft X-Ray Flash Supernova SN 2026gzf
Over a decade earlier, the Dark Energy Camera on the Víctor M Blanco 4-meter Telescope captured images of a blue source in the same position. That was intriguing since it implied that the star was already fairly active before collapsing. DECam later captured several further photographs, revealing that the supernova was becoming increasingly strong. Meanwhile, the NSF-DOE Vera C. Rubin Observatory produced multiband observations of the event in its COSMOS Deep Meiling field, revealing information on the star’s behavior before to the explosion. The DESI on the Nicholas U. Mayall 4-meter Telescope captured a succession of spectra that helped determine the type of supernova and tracked the light as it spread.

Soft X-Ray Flash Supernova SN 2026gzf
Rastinejad’s team employed both Gemini telescopes, as well as the Gemini Multi-Object Spectrographs and the Goodman spectrograph on the SOAR 4.1-meter Telescope, to gain a comprehensive look at the event at all wavelengths. They examined the data and concluded that the star that went supernova was a Wolf-Rayet star. Born with nearly twenty times the mass of the sun, it had already depleted all of its hydrogen and blown off all of its helium in a series of frenzied explosions. The residual core was largely carbon and oxygen, and the mass loss episodes left behind a variety of material shells: a small tight one near to the star that provided the soft X-ray signal, and a larger one further away that created the supernova’s visible brightness.
[Source]

Advertisement

Source link

Continue Reading

Tech

AI startup Hark unveils first product: an affordable, fast computer use agent Hark Handoff

Published

on

Hark, the secretive AI startup founded earlier this year by serial entrepreneur and roboticist Brett Adcock, today announced Handoff, a “computer use agent” (CUA) that it says is among the top-performing in the world at navigating the open web on a user’s behalf — ordering dinner on DoorDash, booking flights on United and Delta, or messaging job candidates on LinkedIn — all autonomously, end-to-end.

Sign-ups open to the public today at hark.com, with availability planned for later this month as part of the initial release of Hark’s software platform.

The company says Handoff recorded the top-ever score on Online-Mind2Web (OM2W), a third-party benchmark with a human-evaluated leaderboard for web agents, posting a 97.7 against 92.8 for OpenAI’s GPT 5.4, 84.1 for Anthropic’s Claude Opus 4.8, and 69 for Google’s Gemini 2.5 Pro.

Hark Handoff benchmark comparison chart

Hark Handoff benchmark comparison chart. Credit: Hark

Advertisement

Hark also says it can serve the model at less than one-tenth the token price of competing frontier models — $0.18 per million input tokens and $2.37 per million output tokens, versus $5 and $30 for GPT 5.5 — with per-turn model latency of 0.8 seconds.

Hark Handoff pricing comparison chart

Hark Handoff pricing comparison chart. Credit: Hark

For each request, Handoff spins up a dedicated virtual computer with its own browser, file system, and terminal, and users can connect existing accounts so the agent can log in and act with their saved addresses, payment methods, and history.

Hark’s research uncovered that despite people spending 75% of their screentime every day in a browser, fewer than 1 in 1000 websites have publicly accessible APIs, making it challenging for AI agents to take over the workload.

Advertisement

In a roughly four-minute produced announcement video posted on YouTube and social media, Adcock — seated in a bare warehouse space that doubles as a metaphor for the company’s build-out — speaks a request aloud to Hark (“let’s liven this place up a bit… let’s do some roses, maybe some cherry blossoms”) and Handoff is shown navigating a florist’s website to place the order, while Adcock narrates that unlike a typical chatbot, Handoff “is always working, it’s looping,” and says he now uses it for “all of my recruiting efforts end to end.” In Hark’s announcement blog post, more demos are shown in realtime and 5x speed.

But big some open questions about Handoff remain, especially for potential enterprise customers and users.

High-scoring benchmarks…but against last generation’s models

Notably, the benchmark comparisons Hark provided to VentureBeat for its Handoff AI agent are against GPT 5.5, GPT 5.4, Opus 4.8, and Gemini 2.5 Pro — the prior generation of frontier models.

Advertisement

The current leaders, OpenAI’s GPT-5.6 and Anthropic’s Opus 5, are absent, as are strong open-source computer-use contenders like DeepSeek V4, Kimi K3, and Qwen3.8-Max.

These newer models haven’t published Online-Mind2Web results, and no third party has posted them to the benchmark’s public leaderboard — meaning Hark’s “top-ever” claim cannot currently be checked against the strongest available systems.

The omission is notable because the newest frontier models have posted their largest gains precisely in computer use: on OSWorld 2.0, a related benchmark covering full computer control, Anthropic’s Opus 5 scores roughly 70.6% versus 55.7% for the Opus 4.8 model Hark chose as its comparison point.

The latency comparison comes with similar caveats: the 6.8-second and 6-second per-turn figures Hark cites for GPT 5.5 and Opus 4.8 were measured by Hark, in Hark’s own harness, with the competing models set to their highest — and slowest — reasoning level. No independent latency measurements exist for comparison.

Advertisement

Asked by VentureBeat whether Hark plans to publish comparisons against those newer models, the company did not specify.

Even within Hark’s own chosen comparisons, the “best” framing has an asterisk: on WebTailBench v2, one of the three benchmarks in Hark’s own results table, GPT 5.5 scores 72.3 to Handoff’s 68.6.

Two of the three benchmarks (WebTailBench and an unnamed internal evaluation) were also run inside Hark’s own harness, with pass rates computed by Hark’s internal LLM judge — conditions the company controls.

Hark’s pricing advantage is far clearer: Anthropic’s newer Opus 5 carries the same $5-per-million-input and $25-per-million-output list price as its predecessor, so Handoff’s roughly tenfold cost savings would hold up even against the current frontier — assuming its benchmark performance does too.

Advertisement

Training and file access

Hark’s research preview describes a sensible-sounding pipeline — supervised fine-tuning followed by asynchronous reinforcement learning using the GRPO algorithm, according to materials shared with VentureBeat prior to today’s announcement — but the company acknowledges it has only done post-training so far, with pre-training “planned for later this year.”

That means Handoff is built on top of a base model Hark did not train. Asked which base model it is, and what mix of proprietary and open data Handoff was trained on, Hark hasn’t yet specified.

Another big question mark for enterprise users: who can access the dedicated virtual computers and the files created on them?

A Hark spokesperson said “security and privacy is a primary focus, but this is a technical preview,” adding the company will share more when the product reaches market at the end of the summer.

Advertisement

Adcock’s history leading up to Hark

Hark is Adcock’s fourth company. He previously co-founded the talent marketplace Vettery (sold in 2018 for roughly $100 million), the air-taxi maker Archer Aviation, and the humanoid robotics unicorn Figure AI.

Hark raised a $700 million Series A round in May 2026 at a $6 billion valuation — led by Parkway Venture Capital, with participation from Nvidia, AMD, Intel Capital, Qualcomm Ventures, Salesforce Ventures, and ARK Invest.

Adcock seeded the company with $100 million of his own money and remains founder and CEO of both Figure and Hark simultaneously, a spokesperson confirmed.

Asked how the two companies interact, the spokesperson said Hark models “are being trained on the Figure robots,” but that Adcock has no plans to combine them.

Advertisement

Adcock’s promotional style has drawn skeptics. In April 2025, Fortune correspondent Jason Del Rey reported that Figure’s much-touted BMW partnership was far more modest than Adcock’s public claims of a robot “fleet” performing “end-to-end operations”: BMW spokesperson Steve Wilson said a single Figure robot was practicing picking up parts during non-production hours.

But the partnership has advanced, and as of June 2026, BMW said the Figure 02 robot supported production of more than 30,000 BMW X3 vehicles during a 10 month-period, and that the next-generation Figure 03 robot was being deployed at the plant for a parts-sequencing role in logistics.

On the social network X, Adcock called the story “mischaracterizations and downright lies” and threatened a defamation suit. Two months later, TechCrunch reported that Adcock skipped a promised live demo at a tech conference and sidestepped questions about the BMW deal onstage.

None of that means Handoff’s numbers are wrong. The agent may well be excellent, and the pricing — if it holds — would undercut every major lab.

Advertisement

Source link

Continue Reading

Tech

Meet the eight startups pitching at Startup Battlefield Australia

Published

on

The applications are in. The TechCrunch Startup Battlefield team made their decisions. Eight Australian startups have earned their spot on the Startup Battlefield stage at Stripe Tour Sydney on August 19 — and they’re ready to pitch for $15,000 in Stripe fee credits, investor attention, and a spot in Startup Battlefield 200 at TechCrunch Disrupt in San Francisco.

They represent the future of Australian tech — from enterprise tech to health and wellness tech to entertainment and media. Here’s who they are.

The eight finalists

Aigentsphere

Independent AI management and governance platform Aigentsphere helps enterprises scale AI safely and well, giving leaders visibility, control, and accountability over their entire AI agent workforce.

Apate.ai

Apate help banks, telcos, and governments dismantle scam operations by engaging scammers in live conversation and turning it into real-time intelligence.

Advertisement

Callease Ai

Callease Ai is the operating system for physical security control rooms, starting with voice AI that automates welfare, escalation, patrol, and many more — the entire calling engine.

Choosey App

Choosey App helps businesses hire instantly and workers get paid instantly through agentic AI-powered matching, workforce management, and same-day pay.

Doomers AI

Doomers AI helps AI and tech companies make product launches trend on X by orchestrating a vetted creator network that reaches the buyers who block ads.

equ

Startup equ helps health and wellness businesses deliver personalized nutrition at scale by embedding equ’s AI engine, built on 10 years of data and 100,000+ users, into their product.

Advertisement

LeadStory

LeadStory is the video intelligence layer powering the next generation of screens, surfacing exact video moments from sports, news, and finance clips to answer natural language queries.

Preve

Preve helps physical therapy clinics improve patient retention and clinical outcomes by automating the treatment plan creation and adherence.


Three will win prizes. One will go to San Francisco. All eight will pitch in front of investors, press, and the Australian tech community. If you want to see the future of Australian startups unfold, register for Stripe Tour Sydney on August 19 →

The stage is set. The judges are ready. Let the pitching begin.

Advertisement

Hosted and MC’d by Isabelle Johannessen, Director of Startup Battlefield.

When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.

Source link

Advertisement
Continue Reading

Trending

Copyright © 2025