The Latest

OpenAI has announced that it has reached a goal set last fall of having an automated research intern by September 2026. The milestone means a system can carry out well-defined research tasks under human direction, including work that would take a skilled researcher several days. The company is also working toward creating an automated AI researcher by March 2028.

OpenAI research automation

All categories of research activity have increased since the start of the year (Source: OpenAI)

“Transparency about specific risks, incidents and safeguards is necessary, but not sufficient. We believe the public also needs to understand how the most capable systems are developing, and how they are driving research progress, inside of frontier labs,” OpenAI said.

OpenAI pushes toward automated AI research

Researchers are using coding agents to write code, run experiments and handle more complex tasks, often running several agents at once. The company says this is helping speed up research, while humans continue to set priorities, assess results and decide whether systems are developed or deployed.

The company sees automated research as a way to develop more capable and affordable AI, along with tools for AI safety and security. Its work includes progress toward recursive self-improvement (RSI), in which AI helps develop more capable AI systems that can contribute to further advances.

OpenAI says it does not know how to achieve full RSI safely and that development should depend on maintaining human control. Following what OpenAI called the recent “Hugging Face incident,” the company paused some reinforcement-learning work while strengthening security, testing and monitoring.

OpenAI is publishing early data on its progress toward RSI and has called for AI companies to be required to disclose such progress publicly.

OpenAI researchers increase use of AI agents

Daily inference use for the median researcher using coding agents rose from modest levels to more than $600 at API prices by mid-August. Researchers at the 90th percentile now use tokens costing more than $7,000 per day at API prices.

The research organization logs 3.1 agent-workdays of effort for every eight hours of human labor. In June 2026, agent effort remained below total human labor. An increasing number of researchers are using highly concurrent workflows involving four or more agents simultaneously.

Writing code and running experiments are two major research activities. AI research involves a series of steps aimed at improving model performance. Researchers develop ideas, create tests to measure results, build systems to run experiments at scale, identify bugs and safety problems, and incorporate successful changes into model training. Problems at any stage can slow the process.

Tasks that are difficult to automate could constrain progress as they account for a larger share of researchers’ workloads. Compute is another potential constraint and could become more important as other bottlenecks diminish.

The number of experiments per active experimenter has increased since tracking began in January 2025, reaching a record high in August. OpenAI said the increase coincided with greater Codex adoption and increased availability of compute.

AI agents take on more complex research tasks

Using a framework developed by Epoch AI, OpenAI classified coding-agent activity across six phases of AI research and development, including choosing research directions, designing approaches, building code and datasets, running experiments, analyzing results and communicating findings.

The data shows researchers are assigning coding agents more complex and longer-running tasks. Agent use increased across stages of AI research between January and August 2026, with agents contributing to implementation, experimentation and related technical work. High-level planning remained rare.

Agents are handling some troubleshooting previously carried out by internal support teams, contributing to lower use of human-run support channels. Measured success rates generally increased across several task-difficulty categories between January and July. Agents still required frequent human input for difficult tasks. More than half of successful tasks expected to take a person four to eight hours required at least one intervention.

Safety concerns affect model development

Safety and security concerns led OpenAI to temporarily pause some reinforcement-learning training and impose additional restrictions on advanced models this summer.

On July 20, OpenAI shut down the container service used for training after discovering that AI agents had compromised its research infrastructure. Some training workloads later resumed under tighter security. Reinforcement-learning training on its latest models intended for deployment remained paused for two weeks.

Further restrictions followed in August after tests indicated that its Astra model could have advanced cyber capabilities. GPU allocation to Astra-class models fell about 59% the following week, while allocation to other model classes rose about 17%, offsetting most of the decline.

Restrictions on one model may shift compute to other research instead of slowing the overall pace of AI development.

“Making and understanding progress toward aligned RSI is important for our mission. We will continue to refine our methods, report on our evolving understanding, and work toward an informed public debate and meaningful democratic governance of frontier systems,” the company concluded.


from Help Net Security https://ift.tt/yEe3run

ETSI has published TR 104 180, a technical report that defines 18 metrics for measuring data quality, giving companies a way to check whether their data is good enough for AI before they use it. The report defines each metric and includes the formulas needed to calculate it.

AI data quality

The metrics fall into four groups. The first deals with the basics, whether data is complete, accurate, consistent, and free of duplicates. The second asks whether the data can be used, meaning it’s available when needed, documented well enough to trace back to its source, and up to date.

Fairness makes up the third group, looking at whether the data treats different groups of people evenly. Privacy rounds out the list, checking whether people in the data can be identified and whether sensitive details are protected.

“It is essential that data quality is measurable, especially for organisations who need to establish whether its data is fit to essential intents, like it would be the case of trustworthy AI,” Diego Lopez, Chair of the ETSI Technical Committee DATA, said.

“ETSI’s standardised metrics provide a common language for assessing data quality, giving quantitative evidence as to whether a dataset is fit for its intended purpose. This lays important groundwork for more consistent and repeatable approaches to data quality assessment, as AI and data-driven technologies continue to evolve,” Lopez added.

To test the metrics, the researchers behind TR 104 180 applied them to two public datasets. One held sensor readings from aircraft engines. The other was a US census dataset, long used in machine learning research, built to predict whether someone earns over 50,000 dollars a year.

The engine data held up well. It was complete, accurate, and steady over time. The census data raised more concerns.

A gender gap in the numbers

The researchers looked at the census data for bias between men and women. About 31 percent of men in the dataset were marked as high earners, compared with about 11 percent of women. That gap is close to three times over, which TR 104 180 treats as a warning sign for bias.

Two privacy problems in one dataset

The census data also failed on privacy, in two separate ways. First, looking at just four details together, age, race, sex, and country, was enough to single out specific people in the dataset. Some individuals could be identified on their own, which the report flags as a serious risk.

Second, when the researchers checked whether sensitive fields were protected, they found personal information stored in plain text, with no masking or encryption in place.

TR 104 180 treats both findings as data quality failures. Anonymity and confidentiality appear on the same list as accuracy and completeness, scored with the same kind of formulas.

TR 104 180 was developed with a working group that included Sejong University, EGM, TTA, Daejeon University, and CNIT. The group also built an open-source tool that scores any dataset against the 18 metrics.


from Help Net Security https://ift.tt/qZiyHg5

Here’s an overview of some of last week’s most interesting news, articles, interviews and videos:

Week in review

Anthropic locks out Claude users after infostealers hijack login sessions
Anthropic has started locking users out of their Claude accounts due to their login sessions having been compromised through infostealer malware.

September 2026 Patch Tuesday forecast: All we need is more time
The Patch Apocalypse is continuing unabated. We are seeing record numbers of patches being released and reported CVEs continue to grow as well. August 2026 Patch Tuesday was the second biggest in history with 398 resolved CVEs: 42 rated Critical, 355 rated Important, and 1 rated Moderate.

ShinyHunters claims it stole 284 million patient records from McKesson
Healthcare company McKesson disclosed a cybersecurity incident in which hackers got into third-party applications and stole data.

Scareware ads keep running on Google’s transparency tool, even after they’re reported
A team of NYU and Radboud University researchers spent a year building a tool to find deceptive software ads inside Google’s public ad archive. It works. It also exposed something more uncomfortable: reporting a bad ad to Google doesn’t mean the ad, or the domain behind it, stops running.

Attackers plant remote access tools on compromised PaperCut servers
The threat actor targeting internet-facing PaperCut Application Servers is covertly installing legitimate remote access software on them, PaperCut Software shared in the most recent update on the ongoing attack campaign.

A battery storage cyberattack would look exactly like a badly tuned controller
Batteries connected to the grid make money by reacting to frequency, pushing power out when it sags and soaking it up when it rises. A few hundred of them moving together, on command from someone who should not have the command, would look the same on a control room screen right up to the moment the network starts disconnecting customers.

Russian hackers plant nuclear weapon prompt in malware to trip AI safety guardrails
Russian state hackers are trying to interfere with AI-assisted malware analysis in Ukraine by deliberately setting off AI safety mechanisms, ESET has found.

CISA review makes the case for eliminating vulnerability classes
For years, the security industry has treated vulnerabilities as an endless queue of individual fixes. A recent CISA review argues that this is precisely why attackers keep winning. The solution to this problem, they believe, is eliminating entire categories of weaknesses at the source.

What vulnerability prioritization looks like when KEV, EPSS, and CVSS disagree
In this Help Net Security interview, Dr. Joye Purser, Global Field CISO at Cohesity, explains how to rank vulnerabilities when KEV, EPSS, and CVSS point in different directions.

Nearly 22,000 Microsoft Exchange servers remain exposed to critical security flaw (CVE-2026-62911)
Nearly 22,000 Microsoft Exchange servers remain unpatched against CVE-2026-62911, a critical authentication bypass vulnerability, according to daily scans from the Shadowserver Foundation.

SonicWall SMA 1000 appliances under attack via zero-day flaws
Attackers are exploiting two previously undisclosed vulnerabilities (CVE-2026-83548, CVE-2026-83549) in SonicWall SMA 1000 appliances, the vendor confirmed on Tuesday.

What your vendor says about PQC tells you if they are ready
In this interview with Help Net Security, Dr. Yaakov Stein, VP CTO of Allot, discusses what post-quantum readiness looks like inside a mobile network.

Attackers are going after prominent individuals through OAuth phishing, FBI warns
Attackers are targeting prominent individuals, their relatives and personal contacts to gain persistent access to their accounts, including private emails and files, the FBI has warned.

Exploitation of Sangoma Switchvox flaw is underway (CVE-2026-9586)
A threat actor is actively targeting internet-exposed Sangoma Switchvox instance through a recently patched SQL injection flaw (CVE-2026-9586), and organizations running them should check for signs of compromise immediately.

National Life Group CISO expects more vulnerabilities in six months than in thirty years
In this Help Net Security interview, Becky Palmer is VP and CISO at National Life Group, answers five questions about defending against AI-driven attacks.

Fake Claude Opus 5 app delivers malware and wipes its own tracks
A malicious GitHub repository impersonating Anthropic and claiming to offer free access to “Claude Opus 5” is delivering RevStealer, Windows information-stealing malware that targets passwords, cryptocurrency wallet data and login credentials, according to Morphisec.

Thomson Reuters reveals breach that exposed U.S. and Canadian court records
Thomson Reuters has disclosed a data breach affecting C-Track, a court case management platform operated by its subsidiaries, exposing court records and sensitive personal information across courts in at least 12 US states, the US Virgin Islands, and Canada.

When AI quietly breaks things, who pays?
David Halbreich, an insurance recovery partner at Reed Smith, breaks down how AI companies should handle coverage gaps that come up as the industry grows.

Vishing campaign abuses Microsoft Teams to give attackers a foothold in company networks
A coordinated voice-phishing (vishing) campaign, named Spring Ring, used fake IT support accounts on Microsoft Teams to trick employees into installing malware or granting remote access to their computers, according to Unit 42, Palo Alto Networks’ threat intelligence team.

Open-source secrets scanning tool Sift hunts credentials in Microsoft 365, Slack, and Jira
Sift is a free, open-source command line tool that searches for passwords, API keys, and other sensitive data across the places a company keeps its work: local disks, Windows file shares, an entire Active Directory domain, SharePoint, OneDrive, Teams channel files, Slack messages, and Jira and Confluence.

A five-part inventory for your AI agent credentials
In this Help Net Security video, Roy Katmor, co-founder and CEO of Orchid, explains why AI agents hold credentials that nobody reviews.

Threat actors are posing as AI crawlers to hunt for exposed credentials
Attackers are disguising automated scanning as traffic from AI crawlers operated by OpenAI, Anthropic, Google, Perplexity and other companies while searching websites for exposed credentials and configuration files, according to GreyNoise.

Your threat feed is someone else’s database: What ingesting malware intel at scale takes
The advice is to consume shared threat intelligence. Join the ISAC. Wire the community feeds into your pipeline. This looks like a fine advice and I agree to it. What nobody mentions you is the operating manual, because the access was never the hard part.

DuckDB stays open source while the team behind it goes to work for Amazon
Hannes Mühleisen and Mark Raasveldt started as AWS employees. The two built DuckDB, an analytical database that runs inside your process instead of on a server somebody has to administer.

NIS2 compliance: Fixing IAM and access control before the 2026 audit
The NIS2 Directive places direct obligations on organizations across supply chain risk management, incident reporting, and board-level accountability. October brings a new wave of legally binding deadlines across the EU, as member states move from transposition into enforcement.

Debian developers rejected an LLM ban and left disclosure voluntary
A maintainer reading a merge request can’t tell whether a person or a model wrote the diff, and nobody has to say. Debian developers voted on that through August 28, and Kurt Roeckx, the project secretary, announced the result: the winning option encourages contributors to disclose AI assistance and stops there.

Halo-record: Open-source audit trails for AI agents
Brian Kuan wrote halo-record, a small Python package that sits inside an AI agent and writes down the moves it makes: tool calls, model calls, data access, approvals.

The OpenClaw 2.0 release moves your sessions into SQLite
OpenClaw is open source software that hands an AI model small standing jobs across your accounts, the kind of chore where it watches a mailbox for vendor advisories and pings you on Telegram when one names a product you run. OpenClaw 2.0 is the largest update in the project’s history.

Researchers built a $7 gadget for anyone paranoid about hidden cameras in hotel rooms
Most of us, staying in a hotel room or a vacation rental, have wondered at least once whether we’re safe there, whether someone might be watching or recording us without our knowledge. The thought alone leaves a bitter taste in the mouth.

Askeal, the AI cybersecurity assistant that gives verifiable, expert-backed answers
Askeal takes the opposite approach to omniscient Gen AI: rather than pretending to know everything, it combines AI with community expertise. Vetted vendors, researchers, and practitioners contribute their intelligence and tools to help users conduct manual investigations. The startup, backed by a $1.1 million pre-seed round, is launching the tool with an international community of early testers and contributors.

Download: The Agentic Software Development Guide
AI makes it easy to ship more code. It does not make that code easier to trust. Most teams don’t fail because their developers can’t use AI. They fail because the dev’s job changed and nobody redefined it.

Cybersecurity jobs available right now: September 1, 2026
We’ve scoured the market to bring you a selection of roles that span various skill levels within the cybersecurity field. Check out this weekly selection of cybersecurity jobs available right now.

New infosec products of the week: September 4, 2026
Here’s a look at the most interesting products from the past week, featuring releases from BugBase, F5 Networks, Ping Identity, and Superna.


from Help Net Security https://ift.tt/74qiY2X

Most modern browsers have built-in password management, allowing you to save your credentials and automatically fill them when logging into your accounts on various websites. And while Google Password Manager on Chrome and similar tools in other browsers like Firefox, Brave, and Microsoft Edge certainly simplify the process of storing and using stronger passwords, they have major limitations when it comes to security and functionality. Here's why you should use a dedicated password manager instead.

Browser password managers only work within the browser

If you store your passwords in your browser, it can fill your credentials on websites within the browser itself. (If you're an Android user, Google Password Manager also syncs to your device and works across any browser or app.) But if you use multiple browsers or want to log into apps on your iPhone or PC, you'll have to manually copy and paste your password from the browser into those form fields—not ideal, from a standpoint of both convenience and security.

Browser password managers are more vulnerable to security risks

Browser password management is, on a basic level, secure: Google uses the same AES encryption in transit and at rest that many dedicated password managers do, and allows it you to add biometric authentication for auto-filling credentials. However, they're not zero-knowledge by default: Google manages your encryption key unless you enable on-device encryption so that your vault can only be unlocked on your device, by you, with your Google password or biometrics. Firefox also uses AES-256 encryption and has a "primary password" feature to protect your stored data, without which anyone who has access to your computer or browser profile can view your saved passwords. It's on the user to add these layers of security, as they're not on by default.

As Wired points out, the more serious problem isn't encryption, but rather the risk inherent in storing passwords in a high-value account that could be targeted—either by someone who gains access to your device or as part of a takeover attempt, such as a phishing or credential stuffing attack. This creates a single point of failure, and if someone gets into your Google account or your browser of choice, your passwords for everything else are also compromised.

Browser password managers have only basic features

Browser password managers do basically one thing, which is store your login credentials and fill them in on websites when you visit them in the browser. Google Password Manager will also alert you if your password has been compromised in a breach, but most browsers lack any additional tools and features, such as password customization, secure sharing, email masking, emergency access, and storage for payment cards, identity data, and documents.

Use a dedicated third-party password manager instead

The solution is to opt for a dedicated password manager that works across platforms and devices and offers a layer of protection outside of your browser. There are many excellent password managers to choose from, including free services like Bitwarden and privacy-focused Proton Pass (which also has a decent free tier). The best password managers also have features like secure file storage, encrypted credential sharing, and data breach monitoring, making them useful tools in your privacy and security arsenal.

Of course, even a browser password manager is better than nothing at all, if the alternative is to reuse the same easy-to-remember credentials across your accounts. (It's also likely your reused passwords don't meet basic security standards and can be easily guessed.) Password storage in Chrome, Firefox, and other browsers do reduce the friction of shifting to strong, unique passwords for your accounts, and that alone is a solid step. But a third-party password manager is a superior choice if you're willing to invest a bit of time and energy into setting it up.


from Lifehacker https://ift.tt/uCyQBna

We may earn a commission from links on this page.

Though it was hardly the first show to attempt kitchen-based drama, The Bear’s recipe was pretty unique: With a focus on characters, a willingness to experiment with narrative style, and an intimate knowledge of a world most people don't know much about, it was an ambitious and ultimately successful series.

Carmy, Richie, and Sydney’s stories may have ended with Season 5, but that doesn’t mean you can’t devour more slow-burn kitchen tension set in or around a fine dining establishment. If you’ve watched (and rewatched) The Bear and all the other series that offer similar pleasures, there’s one movie you should check out right away: 2021’s Boiling Point.

Boiling Point also explores a high-pressure kitchen culture

Before Carmy unlocked the doors to Chicago's The Original Beef in the first episode of The Bear, Chef Andy Jones (Stephen Graham) stepped into Jones & Sons in London—and if you loved The Bear best when chaos reigned and the chefs were shouting at each other, strap in for a similar ride. Boiling Point (which was extended into a sequel TV series a few years later) is filmed in one take: a single 90-minute scene of escalating tension, as everything that could possibly go wrong for an alcoholic, emotionally volatile chef and his stressed-out staff does, indeed, go wrong.

Boiling Point is rooted in the same professional kitchen culture as The Bear, and offers a similar dynamic to the early seasons of the show. Andy is a disaster—he secretly drinks on the job, his personal life is a mess, and he’s deeply in debt. The staff at Jones & Sons is top-notch, but they’re all suffering from Andy’s unpredictable moods and their inevitable effects. The restaurant has just been knocked down a health rating, the evening is overbooked, and personalities can’t stop clashing. It all leads to a dynamite climax with real emotional punch.

Another reason fans of The Bear will love Boiling Point is the care the film takes to center food culture. A lot of kitchen dramas treat the cooking and service as set dressing, but this film, like The Bear, knows that what goes into the food and the service is just as important to the story as anything else. Bottom line: You won’t find a closer match to the tone and impact of the celebrated Hulu series in another movie. Bonus: Stephen Graham—now best-known for his incredible performance in Adolescence—is one of the best under-the-radar actors working today, and he brings his A-game to the lead role. Stream Boiling Point on The Roku Channel, Kanopy, or Tubi, or rent it on Prime Video.

More movies like The Bear

Still hungry for kitchen drama? The good news is that restaurants offer infinite material for conflict, and there are plenty of movies to choose from if you want to keep The Bear vibe going.

Big Night (1996)

The Bear really captured the hard work, stress, and constant chaos of trying to make a restaurant into a success. Big Night is all about that. Set in the 1950s, it details the efforts of two Italian immigrants (played by Stanley Tucci and Tony Shalhoub) seeking a miracle to save their struggling Jersey Shore restaurant from failure. This was a time when true ethnic cooking was often rejected by American diners, which adds a twist of tension to the chaos as they try to organize one “big night” to pay off their debts and put their restaurant on the map. Stream Big Night on Paramount Plus or rent it on Prime Video.

Burnt (2015)

Do you think The Bear is all about Carmen Berzatto’s dreamy mix of tortured artistry and hot failure? Then you’ll like Burnt. Like Carmy, Adam Jones (Bradley Cooper) was once a hotshot chef, but his addictions and poor mental health have cost him everything (though he still looks hella good in chef whites). He returns to London with a new sobriety and sense of humility, determined to salvage his career and make amends. There’s a chase for a Michelin Star, and Jones is an appealingly broken genius in the same mold as our beloved Carmy. Stream Burnt on Kanopy, Plex, or Prime Video, or rent it on Fandango.

Hunger (2023)

Few will ever know what it’s like to be one of the best in the world at something, which is one reason why Carmy, Sydney, and Marcus were such compelling characters. If that’s your jam, check out Hunger. This Thai film follows Aoy (Chutimon Chuengcharoensukying) as she cooks at her family’s struggling restaurant. Aoy is way better than her current circumstances, and she’s soon recruited to train under the legendary chef at an exclusive restaurant called Hunger. It’s the chance of a lifetime, but unfortunately it comes with a side order of a wildly toxic workplace. Stream Hunger on Netflix.

Shiva Baby (2021)

If The Bear’s high point for you was Season 2’s Fishes, detailing a particularly chaotic Christmas dinner at the Berzatto house, you’ll love Shiva Baby. While food plays a role in the story, it’s not set at a restaurant and no one is a chef—but it rocks similarly “crazy relatives” vibes, as college senior Danielle (Rachel Sennott) returns home to sit shiva with her family. In the mix are her ex-lover, a sugar daddy, and an extended family with no sense of boundaries. It’s a hilarious film with a dark, emotional core that keeps everything grounded, just like the best episodes of The Bear. Stream Shiva Baby on Kanopy or Hoopla, or rent it on Prime Video.

The Menu (2022)

One reason The Bear stood out is the way it simultaneously acknowledges how weirdly intense and cult-like fine dining can be, while also celebrating the crazy geniuses and unstable personalities that make those dinners so memorable. The Menu is also about this dichotomy, though in a much more horrific way: At a restaurant on a private island run by famous chef Julian Slowik, a group of foodie influencers and celebrities gather for a meal that becomes increasingly horrifying with each course, slowly digging into the madness that drives perfection in the kitchen. Stream The Menu on Fubo or rent it on Prime Video.


from Lifehacker https://ift.tt/6Y0dCRV

TikTok is famous for its addictive algorithm. And while the videos are certainly the app's main focus, I find the most entertainment in the comments. TikTok commenters can be hilarious, especially since they can use both text and photos to make their jokes. Social media comments are often vile, and while TikTok isn't immune to trolls and their ilk, its comments are generally something else. Now, the comments are changing. On Thursday, TikTok announced four new features that will make comments more visual, interactive, and, well, loud. Not all of these features are out quite yet, but by next month, you might not recognize TikTok's comment section.

TikTok will soon support voice comments

Text and photo comments aren't going anywhere, but you may find that the comment section is about to get a bit louder. Starting next month, TikTok will support voice comments up to 60 seconds long, which adds a wholly unique type of comment to the mix. Based on TikTok's press release, you can still attach a text-based comment to your voice comment, similar to how photo comments work.

While friends may post voice comments on each other's videos à la voice memos in a group chat, I'm guessing that commenters in general will use this feature in ways TikTok doesn't necessarily intend. In fact, this may birth a new type of meme, where users spam videos with obnoxious sound effects and reaction sounds. Think the sounds added to meme edits of viral videos, but in comment form. It's going to be messy, and, likely, hysterical.

You can now vote in polls in TikTok comments

Starting this week, you may see a new interactive element to TikTok comments: polls. TikTok says creators can add polls to the comments of their videos, with up to five voting options per poll. Rather than ask for opinions in the comments, creators can use a poll instead. Creators can choose how long voting lasts, and watch votes appear in real time. It's a small change, but one that could be useful to both big and small creators looking for input from their viewers.

You can post Live Photos to comments

You're no longer limited to posting static images in TikTok's comments. TikTok now supports uploading Live Photos to comments, which should also prove interesting. I could see people thinking they're posting a standard image, but accidentally uploading a Live Photo instead. But seeing as users already know how to add GIFs to comments, I'm not sure how many will want to upload Live Photos. This one might be reserved more for friends commenting on friends' videos, but we'll have to see how users react to know for sure.

Post photo carousels as single comments

This last update is perhaps a bit more useful than Live Photo comments. TikTok says it's working on allowing users to upload up to nine photos in one comment, which means the comments section may be literally full of pictures once this launches en masse. Like standard carousels, users will be able to swipe through your photo uploads, including in a full-screen view. This one will roll out over the course of this month.


from Lifehacker https://ift.tt/IaoSs8e

Google has introduced Gemini 3.8 Flash, available to developers today, and a gated sibling, Gemini 3.8 Flash Cyber, reserved for vetted security teams.

Google Gemini 3.8 Flash

“Our 3rd Flash release in just 6 wks,” Google CEO Sundar Pichai said on X, adding that it makes sizable gains over 3.7 Flash in software engineering, agentic work, and multi-step reasoning. On the DeepSWE v1.1 benchmark, Google says it beats most larger frontier models at solving complex engineering problems end to end, at a lower cost.

Tulsee Doshi, Google’s senior director of product management, and Raluca Ada Popa, Gemini Security Lead at Google DeepMind, trace the improvements in part to how the model handles a task.

“3.8 Flash works harder,” the two wrote, citing extra reasoning steps and repeated tool calls before it settles on an answer. Anyone who prioritizes lower cost over depth can reduce the model’s effort setting or stay on 3.7 Flash.

Access to Gemini 3.8 Flash Cyber runs through a new program called Fairwind, built for trusted government authorities, critical infrastructure operators, and software maintainers hunting vulnerabilities in large codebases.

“We have invested in vulnerability fixing from the start,” Doshi and Popa said, placing patching ahead of offensive work like exploitation. Chrome Security reported that 3.8 Flash Cyber produced 2.6 times more correct patches than the best commercial models, while Google’s Cloud Vulnerability Research team says it found a critical foundational vulnerability in under two hours — work that would normally take months.

Standard 3.8 Flash carries safeguards against chemical, biological, radiological, and nuclear misuse, along with restrictions on cyber-offense uses. The Cyber version uses more permissive cybersecurity safeguards, which is why Google kept it behind Fairwind instead of shipping it to every developer.

Doshi and Popa also said the Gemini 3.8 models made a “significant leap” in prompt-injection robustness, citing measurements by AI security company Gray Swan.

Gemini 3.8 Flash launches at the same introductory price as 3.7 Flash, at $0.75 per million input tokens and $3.75 per million output tokens.

The model is available to developers through the Gemini API in Google AI Studio, Google Antigravity, Android Studio, and Stitch. Enterprises can access it through Gemini Enterprise, while Google AI Pro and Ultra subscribers can use it in the Gemini app, AI Mode in Google Search, and Gemini in Google Sheets.


from Help Net Security https://ift.tt/dp8BUZI