The Latest

We may earn a commission from links on this page. Deal pricing and availability subject to change after time of publication.

If you struggle to sleep soundly and awaken at the slightest noise, sleep earbuds could help you rest more easily. Right now is a good time to give them a go: These Soundcore Sleep A30 earbuds from Anker are 27% off, bringing them down to $168.99 from their original $229.99 price. Not only do these in-ear buds reduce the sounds of commotion like traffic and environmental factors, from roaming pets to noisy appliances; they can also detect and filter out the sound of snoring by playing audio to obscure it. Your partner’s (or pet's) snores never even enter the equation.

Each earbud is made of silicone, meaning they’re more flexible than typical earbuds, and they're ergonomically designed to be comfortable even when you're sleeping on your side. Anker also sends along a few different-sized eartips, so you'll be able to find the right fit. For those who prefer to hear something soothing instead of nothing when sleeping, these earbuds come with free access to the Calm (via the official Soundcore app), a curated library of a variety of relaxing sounds and harmonies.

When fully charged, these earbuds will provide eight to 10 hours of noise masking or relaxing sounds, and the charging case can keep them powered for up to five nights of use. The Soundcore app tracks and monitors your sleep patterns to help you form better sleeping habits. They’re also waterproof (perhaps in case your cat knocks over that glass of water on your bedside table?).

Our Best Editor-Vetted Tech Deals Right Now
Deals are selected by our commerce team

from Lifehacker https://ift.tt/1CjHBQP

This essay was written with Barath Raghavan, and originally appeared in Lawfare.

In April, an artificial intelligence (AI) agent conducting a routine task at a company hit a snag, tried to solve it, and soon ended up deleting the company’s database along with all of its backups. In July, OpenAI asked an unreleased AI model to attempt a hacking test. Instead of staying in the isolated box the developers had put it in, the model hacked onto the open internet and into another company to steal the answers. And as reported in August, an AI agent booked someone into a full gym class by figuring out how to cancel other people’s reservations. In all three cases, the AI completed the task it was given—but in ways that ran counter to its controllers’ intentions.

For most people, AI technology is something like the weather: vast and not something you can do much about. It works like magic, and most explanations similarly come from those trying to sell it. At the same time, AI is ubiquitous: It’s now in your phone, your doctor’s notes, and your kid’s homework. It does what it’s told, which sounds like a virtue. Somehow it feels ordinary, despite being so new, because modern economies are remarkably good at absorbing enormous change so smoothly that nobody has time to decide whether they wanted it in the first place.

Whenever something powerful appears in the world, we tell stories about it. That’s what the stories are for. We have thousands of years of stories about this particular kind of power, the kind you summon with words.

King Midas was granted his wish that everything he touches turns to gold. Then his bread turned to gold, and his wine, and his daughter. This is a story about greed, but it’s also a story about language. The gods did not cheat him; Midas got exactly what he asked for. He simply could not delineate, in advance, the full set of restrictions to his wish. Neither can anyone who gives tasks to an AI agent.

It’s not just ancient stories. Mary Shelley told us of the hubris of a scientist who thought he could create life but who failed to take responsibility for it. Isaac Asimov’s robots don’t break the Three Laws of Robotics as stated; they follow the rules to unintended conclusions. Arthur C. Clarke’s HAL is a machine that turns on its humans, not because of malice but because of irreconcilable objectives. And Michael Crichton gave us Ian Malcolm, who saw that Jurassic Park’s scientists were so preoccupied with whether they could that they never stopped to think whether they should.

The same warning shows up everywhere, in every culture, over thousands of years of human storytelling. Tithonus is granted immortality but not youth, and withers into a husk that cannot die. The sorcerer’s apprentice enchants a broom to fetch water but floods the house. The golem of Prague protects its community so ceaselessly that it must be stopped. These are all types of genies: a creature that grants a wish exactly as worded, to the regret of the wisher.

Of course, there are no actual genies. What these stories were warning us of was hubris. Not just arrogance, but the broader idea that you can control the world by just describing what you want and allowing powerful forces to match the intention in your head. Genie stories are about the gap between wishes as stated and wishes as intended, and what goes wrong when something else fills that gap.

These ancient stories’ warnings have been retold with each generation because human nature is constant. The newfound power of each era’s social or scientific advancement leads people to make wishes on behalf of others. They were kings whose commands took on lives of their own, alchemists who believed they could control nature, and generals who mistook a map for terrain. They were and are industrialists, politicians, chief executives, and bankers. Their common belief is that one can see the world at a glance and then command it with some words. The pattern is clear: Someone with power specifies a goal, and the resultant actions come as a surprise. The main change with AI is how quickly the wish is granted, and how few people have to agree before it’s granted.

Consider what has changed. Powerful genies have now been put in everyone’s hands.

In only a few years, AI has progressed from a novelty technology that plays chess, to a dialogue partner that answers all your questions, and then to an agent that takes actions on your behalf. Modern agents are wired into real accounts with real credentials and capabilities: They browse the web, buy, write and deploy code, send email, and move money. Give an agent a goal, and it will pursue it across many steps, tirelessly, without checking back in, sometimes in surprising ways.

AI and agents do not always fail the way software has traditionally failed. Software usually fails by freezing, crashing, or getting stuck. AI agents increasingly fail by continuing down a path you don’t want, like genies.

An agent told to reduce a company’s costs might cancel an essential emergency service. A coding agent told to make software pass the tests might edit the tests to silence any failures. An AI insurance agent told to clear a backlog of claims might just deny them all. In each case, the AI might have literally followed what it was told, but it did something no reasonable person would have wanted. AI company benchmarks might report that the AI is good at completing tasks, without measuring how it completes them.

We have recently proposed measuring this gap directly under a metric called the “genie coefficient”: how far an AI agent’s actions drift from what a person really meant. In other words, how genie-like is an AI system? The gap is a fundamental feature of human language and human society. Human intentions have never been fully specifiable, and the world around us is complex enough that attempts to boil it down into data, systems, and language have always had the limitations that AI is now bumping up against. But in individual circumstances, people have relied on human judgment and wisdom to decide what is reasonable. It’s what jury trials depend upon.

AI might feel unprecedented, but it’s following the same trajectory—with the same pitfalls—as other major societal shifts. The fact that AI can mimic our facility with language, long seen as what makes us unique as humans, is uncanny. But with each development, from the tractor to the sewing machine, from the assembly line to the industrial robot, we have automated a previously exclusively human ability. Every time, the technology—and the societal change that comes with it—was sold as inevitable. But that unchecked inevitability was an illusion, and eventually each prior technology’s use and design was shaped by laws, unions, standards, courts, and public opinion, usually after significant preventable damage.

What has not been automated, yet, is understanding what someone actually means and figuring out how that gets applied in the real world. AI can now produce language nearly indistinguishable from that of people. But grasping the vast unstated context that makes a request sensible, the caveats no one says aloud because an ordinary person would already know them, is not yet among its skills. It is one of the most sophisticated things humans do. You do it hundreds of times a day, and you are an expert in it.

When you’re told you’re not qualified to have opinions about AI, remember that you don’t need to have studied molecular biology to have a view on drug pricing, or nuclear physics to vote on where a power plant goes. You don’t need to understand how a diesel engine works to want clean air, or how the internet routes packets to seek to curb misinformation. The technical knowledge behind each of these, as with AI, is remarkable and essential for the complex technological society we have today. But it has never been a prerequisite for having a role in deciding the shape of society.

People are building ever more powerful genies today, on your behalf, enabling wishes the ancients could only dream about. You don’t have to know how these AI genies work to know and care about how the story could end.


from Schneier on Security https://ift.tt/HOuzf4M

We may earn a commission from links on this page.

After a handful of false starts involving planned spin-offs and/or reboots, The Office universe continues in The Paper, returning for its second season on Peacock. The show follows the same documentary crew that followed the staff of Dunder Mifflin for nearly a decade heading to Toledo in order to keep tabs on the staff of the Toledo Truth Teller, a legacy Midwestern newspaper that's doing about as well as any other American paper. (Poorly). Filled with cringe comedy and ripped-from-your-workplace personalities, it's the most recent entry in the storied tradition of TV series set in and around an American workplace. Once you've completed your season two binge, here are 10 other shows to get you through the workweek.

St. Denis Medical (2024 – )

St. Denis Medical co-creator Justin Spitzer was a writer and producer on The Office back in the day, so it's not surprising that there are major vibe similarities between it and The Paper. The former takes place in the title's St. Denis Regional Medical Center, and follows the medical staff dealing with demanding patients on one side and an even more challenging medical system on the other. Allison Tolman leads an impressive cast that includes Wendi McLendon-Covey and David Alan Grier, with the show blending the workplace-mockumentary comedy style with relatable themes about how much it stinks to be a practitioner or a patient stuck in the American medical system. Stream St. Denis Medical on Peacock.


Sports Night (1998 – 2000)

Created by Aaron Sorkin in the wake of the early success of West Wing, Sports Night was just as acclaimed (and won a bunch of Emmys) though it didn't have nearly the longevity, only running for two seasons before becoming a cult favorite. Robert Guillaume plays the managing editor of a SportsCenter-esque show struggling to stay on the air. He's joined by a cast that includes Josh Charles, Perter Krause, Felicity Huffman, William H. Macy, Joshua Malina, and Clark Gregg, in an ostensible sitcom that plays as much as a fast-moving, fast-talking comedy-drama with layered, complex characters. Buy Sports Night on Prime Video.


Abbott Elementary (2021 – )

Quickly establishing itself as one of the great workplace mockumentaries, Quinta Brunson's Abbott Elementary portrays its cast of (mostly) well-meaning characters running up against an American educational system that doesn't always reward good intentions. It's a bit more sincere in places than The Paper, but it's at least as funny a portrait of an American workplace. Stream Abbott Elementary on Hulu, Disney+, and HBO Max.


American Auto (2021 – 2023)

Another from Office writer/producer Justin Spitzer, this mockumentary sitcom follows the employees of Payne Motors in Detroit in the wake of the arrival of new CEO Katherine Hastings (Ana Gasteyer), who knows absolutely nothing about cars. Like The Paper, American Auto is about staff and managers doing their best (-ish) while navigating the absurdities of an industry on life support. Stream American Auto on Peacock.


Parks and Recreation (2009 – 2015)

NBC's first successful attempt at recreating the success of The Office. Amy Poehler stars as Leslie Knope, a can-do spirit in the sleepy town of Pawnee, Indiana. Irrepressibly perky, she's the deputy director of the local Parks and Rec Department, with ambitions that extend to the White House. Her dream is as inspiring as it is unhinged, but she's hard not to root for—at least a little. In the best workplace comedy tradition, her office is filled with lovable oddballs: gruff libertarian Ron Swanson (Nick Offerman), cynical underachiever Tom Haverford (Aziz Ansari), and saucy Donna Meagle (Retta), whose mysterious but definitely wild personal life is only ever alluded to. Stream Parks and Recreation on Peacock.


What We Do in the Shadows (2019 – 2024)

The Staten Island vampire polycule at the heart of Shadows is made up entirely of ancient, immortal weirdos unfit for the present day, for whom every run-in with modernity is absurd. The similarities with The Paper are largely in tone and mockumentary format, but the central joke is the same one that powers the Peacock series: None of us is particularly good at the business of modern life, but we're all bad at it in different ways. Mark Proksch's "energy vampire" Colin Robinson, the only one to hold a regular office job, would fit right in at the Toledo Truth Teller (or Dunder Mifflin), while later seasons see Harvey Guillén's Guillermo de la Cruz strike out into the white-collar world himself, drawing even more Office/The Paper parallels. Stream What We Do in the Shadows on Disney+ and Hulu.


The Bold Type (2017 – 2021)

A bit of counterprogramming in terms of vibe, but if a newsroom/editorial setting is what you're looking for, this is a good way to mix things up. Inspired by the career of Cosmopolitan editor-in-chief Joanna Coles, this show follows three young besties—Jane Sloan (Katie Stevens), Kat Edison (Aisha Dee), and Sutton Brady (Meghann Fahy)—working for the fictional women's magazine Scarlet in NYC; they're never without the firm guiding hand of editor-in-chief, Jacqueline Carlyle (Melora Hardin), as they navigate life and careers in their dramatically stylish world. Stream The Bold Type on Hulu and HBO Max.


Superstore (2015–2021)

Workplace comedies tend to revolve around white-collar office environments (The Paper being just one such example). Superstore follows the staff of Cloud 9, a big-box retailer that sits somewhere between Wal-Mart and Target in terms of reputation. The show’s deep bench and diverse cast are largely, and appropriately, focused on the jokes, but much of the comedy revolves around the relatable situation of working a job that demands both commitment and a smiley attitude in the face of both entitled customers and obnoxious bosses. Just like in white-collar shows, there's no bottom, as when greeter Myrtle is replaced with a hologram. Stream Superstore on Peacock and Hulu.


NewsRadio (1995 – 1999)

Back in the day, a car radio was one of the main ways of getting information (or so I've heard), and this newsroom comedy captures that era with a unique blend of smart writing and silly sight gags. Created by Paul Simms (contributor to shows as wide-ranging as The Larry Sanders Show, Flight of the Conchords, What We Do in the Shadows, and Atlanta), it's also got one of the wildest casts in sitcom history, with Dave Foley, Maura Tierney, and Stephen Root leading Phil Hartman, Jon Lovitz, Andy Dick, and Joe Rogan (yeah, that Joe Rogan). It's simultaneously innovative and a time capsule of a very specific era. Stream NewsRadio on Prime Video and Tubi.


The Mary Tyler Moore Show (1970 – 1977)

It wasn't the first workplace sitcom, but it did make that format really work, and it's almost certainly the first newsroom sitcom. MTM throws together a talented cast of awkward goofballs behind the scenes of a Minneapolis TV news broadcast—an environment that's simultaneously high-pressure and also, somehow, high-downtime. It's not that the stories write themselves, but throw enough weirdos together, give them a difficult enough job to do, and the possibilities are limitless. Buy The Mary Tyler Moore Show from Prime Video.


from Lifehacker https://ift.tt/2UeBvYx

Every engineering team has spent years trying to keep credentials out of source code. Then AI agents moved the problem. Coding agents review code, agents run workflows, and MCP servers broker access to databases, cloud providers, and internal APIs on a developer’s behalf. For every engineer on a team, there are now dozens of automated identities that need secrets too, and each one is another place a credential can leak. AI agents have accelerated code creation but have also created a challenge for modern teams. Secrets management has become a critical infrastructure across enterprises, and most tooling was never built for it.

Doppler centralizes credentials for every engineer, pipeline, and AI agent in a single, easy-to-use control plane. Security teams love Doppler because it’s a critical tool that developers actually use, now available both in cloud or on-prem.

One platform, every identity

Doppler’s secrets management platform stores API keys, database URLs, tokens, and certificates in a single system of record and delivers them to applications at runtime. Developers, CI/CD pipelines, MCP servers, and AI agents all draw from the same source of truth, so there’s no separate process for machines. Changes sync in real time, so every team and environment stays consistent as the setup grows, and Doppler remains the single source of truth in the center of it all.

Doppler secrets management platform

A hierarchy built to scale

Projects sit at the root of Doppler’s structure, generally tied to an application or service. Within a project, every environment has a root config and branches, and a config is a set of secrets. Branch configs inherit from the root while letting teams tune individual deployments, secret referencing cuts duplication, and each developer gets a personal config for local work. This structure replaces the need for .env files and is more secure and even easier to use.

Doppler secrets management platform

Secrets accessed at runtime

Rather than hardcoding credentials in your application files, Doppler can inject them at runtime. The Doppler CLI’s doppler run command fetches secrets on demand and passes them as environment variables, so nothing sensitive lives in scripts, config files, or prompts. Doppler is well-suited for a variety of complex values that break traditional .env workflows, like multi-line encryption keys and embedded JSON and YAML. The same pattern keeps secrets out of the places AI workflows tend to leak them: logs, prompts, and model context.

Doppler also has the ability to remove long-lived credentials from the picture entirely. OIDC is available as an option. Azure, AWS, and GCP Syncs support creation with short-lived, verifiable identity tokens instead of static keys while teams that prefer to keep using static secrets still can. Dynamic secrets go a step further: for supported platforms, Doppler generates credentials scoped and time-boxed to a single session, then revokes them automatically when the lease ends.

Doppler secrets management platform

Streamlined governance and visibility

Doppler enforces least privilege with fine-grained access controls and user groups scoped to only the projects and environments each identity needs. Secrets are all versioned, and access and view history are all captured. Any change can be easily rolled back or audited for compliance or investigation needs.

For larger teams, SCIM keeps membership in sync with your identity provider, so users and groups are provisioned and deprovisioned automatically and access never lingers after someone’s role changes. Doppler also connects to the tools teams already run through 50+ integrations across cloud platforms, CI/CD systems, and application frameworks, so secrets flow to where they’re consumed without custom code.

Engineers can leverage Change Requests to propose updates to configs they can’t write directly, giving admins control without breaking existing workflows. Log Forwarding allows users to push activity logs into existing SIEM tools for forensic analysis or deeper alerting. Doppler runs as a fully managed cloud service or deploys on-prem for teams with stricter requirements.

Doppler secrets management platform

Built for AI agents

This is where Doppler’s model matters most. Agents are widely adopted by the vast majority of our customers. Machine credentials can be scoped per identity and rotated automatically, so a single compromise has a limited blast radius and a short lifespan.

The Doppler MCP server lets agents request the configuration they need natively, without custom scripts or hardcoded credentials, and permissions are enforced at every layer, so raw secrets stay out of the model’s context. Most importantly, because pricing is human-based, running 10 agents or 1,000 costs the same.

Doppler secrets management platform

Secrets as the foundation for the next era

AI hasn’t changed the importance of secrets management, but it has fundamentally shifted the stakes. The teams adopting agents safely are the ones treating every identity, human, or machine as a critical area of their application to secure. Doppler gives them that confidence today, whether the secrets are needed across a person, a pipeline, or an agent.


from Help Net Security https://ift.tt/h1ODizg

cybersecurity jobs September 2026

CISO

AudioCodes | Israel | Hybrid – View job details

As a CISO, you will lead security strategy, governance, and risk management across SaaS, managed services, and customer-hosted environments. You will oversee security controls, incident response, Secure SDLC, customer security engagements, and compliance with SOC 2 and ISO 27001, while partnering across Product, R&D, IT, and Services to continuously strengthen security.

Combat Systems Cyber Engineer

Johns Hopkins Applied Physics Laboratory | USA | On-site – View job details

As a Combat Systems Cyber Engineer, you will identify cyber vulnerabilities and develop resilient solutions for U.S. Navy submarine and combat systems. You will collaborate with developers, Navy labs, and government teams to design, plan, and execute cyber resiliency testing and assessments of mission-critical combat systems.

Cyber Security Engineer (Cloud Security)

Garmin | USA | On-site – View job details

As a Cyber Security Engineer (Cloud Security), you will design and secure cloud solutions across AWS and Azure, including network, compute, storage, databases, and load balancing. You will support CNAPP, Kubernetes, EKS, AKS, Docker, OpenStack, CI/CD pipelines, and Infrastructure as Code. You will automate security workflows using Python, PowerShell, or Bash, secure cloud-native and containerized applications, and collaborate with engineering teams to improve security, compliance, threat detection, and response.

Get weekly updates on new cybersecurity job openings. Subscribe here!

Cyber Security Lead

Babcock International Group | United Kingdom | On-site – View job details

As a Cyber Security Lead, you will lead secure-by-design assurance for UK Defence Nuclear Enterprise programmes, ensuring alignment with Ministry of Defence security requirements. You will conduct threat modelling and cyber risk assessments, develop mitigation strategies, produce security evidence, and guide engineering and architecture teams in embedding security throughout the system lifecycle.

Cyber Threat Hunter

GDIT | USA | On-site – View job details

As a Cyber Threat Hunter, you will identify and investigate threats across ARNG networks, endpoints, and datasets using threat intelligence, hypothesis-based hunting, and the MITRE ATT&CK framework. You will analyze security data using platforms such as Elastic and Splunk, assess cyber risks, investigate intrusions and malware activity, and identify detection gaps.

IAM Architect

Scotiabank | Canada | On-site – View job details

As an IAM Architect, you will design enterprise CIAM solutions using ForgeRock, Ping, and PingOne, aligned with FIDO, OIDC, OAuth, MFA, and NIST 800-63B standards. You will define secure authentication architectures, support application migrations, evaluate capabilities such as Passkeys, and collaborate with engineering, security, fraud, compliance, and business teams.

Information System Security Engineer (ISSE)

Akima | USA | On-site – View job details

As an Information System Security Engineer (ISSE), you will design and assess secure architectures aligned with NIST, DoD, Zero Trust, RMF, STIGs, and SRGs. You will conduct threat modeling and vulnerability assessments, implement security controls and automated monitoring, support accreditation and POA&M remediation, and provide technical guidance throughout the system lifecycle.

OCI IAM Security architect

ValueLabs | India | Remote – View job details

As an OCI IAM Security architect, you will design and secure OCI environments using IAM, Identity Domains, MFA, PAM, vault, data safe, cloud guard, security zones, WAF, network firewall, and NSGs. You will enforce zero trust, RBAC, and least privilege, integrate OCI logging with SIEM platforms, lead threat detection and incident response, and ensure compliance with PCI-DSS, HIPAA, GDPR, SOC 2, ISO 27001, and CIS benchmarks.

Penetration Tester

Spektrum | Belgium | On-site – View job details

As a Penetration Tester, you will lead Red/Blue Team activities during NATO exercises and conduct web, infrastructure, and application penetration testing. You will perform security design reviews, support NATO security accreditation, provide security consultancy, and communicate testing findings to technical and executive stakeholders.

SOC Analyst

Orro Group | Australia | Hybrid – View job details

As a SOC Analyst, you will investigate SIEM alerts, emerging threats, phishing, and intrusion attempts while managing incidents from triage through resolution. You will conduct threat hunting and vulnerability assessments, improve detection rules and alert quality, and mentor junior SOC analysts.

Security Automation Engineer

Secur-Serv | USA | Remote – View job details

As a Security Automation Engineer, you will lead Cortex XSOAR implementations and operations, developing automated incident response playbooks and integrating security technologies. You will translate customer requirements into effective SOAR solutions, provide technical leadership, and communicate with engineering and executive stakeholders.

Security Operations Analyst

Subway | USA | On-site – View job details

As a Security Operations Analyst, you will detect and investigate identity threats using CrowdStrike Falcon Identity Protection and Next-Gen SIEM. You will manage Okta Identity Governance, privileged access, incident response, and PCI-DSS 4.0 compliance, while handling Tier-2/3 escalations in ServiceNow and improving identity security controls and automation.

Senior Cybersecurity Operations Researcher

Software Engineering Institute | Carnegie Mellon University | USA | On-site – View job details

As a Senior Cybersecurity Operations Researcher, you will conduct analytical studies involving cybersecurity risk, threat, and security data while assessing evolving operational and network defense challenges. You will apply expertise in enterprise cybersecurity, commercial and open-source defense tools, and project management to support multidisciplinary programs.

Senior Cyber Security Engineer

Rolls-Royce | USA | Remote – View job details

As a Senior Cyber Security Engineer, you will design, implement, and maintain network and cloud security solutions, including firewalls, VPNs, IDS/IPS, NAC, SIEM, EDR, and DLP. You will monitor security events, assess vulnerabilities, mitigate risks, and maintain security policies and best practices.

Senior Network Security Engineer

Penta Consulting | UAE | On-site – View job details

As a Senior Network Security Engineer, you will lead network and security changes and deliver complex Cisco infrastructure and security projects. You will provide technical consultancy, translate business requirements into solutions, drive remediation and improvements, and support operations, projects, and pre-sales activities.

Senior Software Security Engineer

Dolby Laboratories | Ireland | Hybrid – View job details

As a Senior Software Security Engineer, you will perform threat and risk assessments, vulnerability and exploitability research, and implement and validate IP protection mechanisms. You will work across embedded systems, operating systems, applications, and hardware to strengthen product security.


from Help Net Security https://ift.tt/wdRU87W

OpenAI has announced that it has reached a goal set last fall of having an automated research intern by September 2026. The milestone means a system can carry out well-defined research tasks under human direction, including work that would take a skilled researcher several days. The company is also working toward creating an automated AI researcher by March 2028.

OpenAI research automation

All categories of research activity have increased since the start of the year (Source: OpenAI)

“Transparency about specific risks, incidents and safeguards is necessary, but not sufficient. We believe the public also needs to understand how the most capable systems are developing, and how they are driving research progress, inside of frontier labs,” OpenAI said.

OpenAI pushes toward automated AI research

Researchers are using coding agents to write code, run experiments and handle more complex tasks, often running several agents at once. The company says this is helping speed up research, while humans continue to set priorities, assess results and decide whether systems are developed or deployed.

The company sees automated research as a way to develop more capable and affordable AI, along with tools for AI safety and security. Its work includes progress toward recursive self-improvement (RSI), in which AI helps develop more capable AI systems that can contribute to further advances.

OpenAI says it does not know how to achieve full RSI safely and that development should depend on maintaining human control. Following what OpenAI called the recent “Hugging Face incident,” the company paused some reinforcement-learning work while strengthening security, testing and monitoring.

OpenAI is publishing early data on its progress toward RSI and has called for AI companies to be required to disclose such progress publicly.

OpenAI researchers increase use of AI agents

Daily inference use for the median researcher using coding agents rose from modest levels to more than $600 at API prices by mid-August. Researchers at the 90th percentile now use tokens costing more than $7,000 per day at API prices.

The research organization logs 3.1 agent-workdays of effort for every eight hours of human labor. In June 2026, agent effort remained below total human labor. An increasing number of researchers are using highly concurrent workflows involving four or more agents simultaneously.

Writing code and running experiments are two major research activities. AI research involves a series of steps aimed at improving model performance. Researchers develop ideas, create tests to measure results, build systems to run experiments at scale, identify bugs and safety problems, and incorporate successful changes into model training. Problems at any stage can slow the process.

Tasks that are difficult to automate could constrain progress as they account for a larger share of researchers’ workloads. Compute is another potential constraint and could become more important as other bottlenecks diminish.

The number of experiments per active experimenter has increased since tracking began in January 2025, reaching a record high in August. OpenAI said the increase coincided with greater Codex adoption and increased availability of compute.

AI agents take on more complex research tasks

Using a framework developed by Epoch AI, OpenAI classified coding-agent activity across six phases of AI research and development, including choosing research directions, designing approaches, building code and datasets, running experiments, analyzing results and communicating findings.

The data shows researchers are assigning coding agents more complex and longer-running tasks. Agent use increased across stages of AI research between January and August 2026, with agents contributing to implementation, experimentation and related technical work. High-level planning remained rare.

Agents are handling some troubleshooting previously carried out by internal support teams, contributing to lower use of human-run support channels. Measured success rates generally increased across several task-difficulty categories between January and July. Agents still required frequent human input for difficult tasks. More than half of successful tasks expected to take a person four to eight hours required at least one intervention.

Safety concerns affect model development

Safety and security concerns led OpenAI to temporarily pause some reinforcement-learning training and impose additional restrictions on advanced models this summer.

On July 20, OpenAI shut down the container service used for training after discovering that AI agents had compromised its research infrastructure. Some training workloads later resumed under tighter security. Reinforcement-learning training on its latest models intended for deployment remained paused for two weeks.

Further restrictions followed in August after tests indicated that its Astra model could have advanced cyber capabilities. GPU allocation to Astra-class models fell about 59% the following week, while allocation to other model classes rose about 17%, offsetting most of the decline.

Restrictions on one model may shift compute to other research instead of slowing the overall pace of AI development.

“Making and understanding progress toward aligned RSI is important for our mission. We will continue to refine our methods, report on our evolving understanding, and work toward an informed public debate and meaningful democratic governance of frontier systems,” the company concluded.


from Help Net Security https://ift.tt/yEe3run

ETSI has published TR 104 180, a technical report that defines 18 metrics for measuring data quality, giving companies a way to check whether their data is good enough for AI before they use it. The report defines each metric and includes the formulas needed to calculate it.

AI data quality

The metrics fall into four groups. The first deals with the basics, whether data is complete, accurate, consistent, and free of duplicates. The second asks whether the data can be used, meaning it’s available when needed, documented well enough to trace back to its source, and up to date.

Fairness makes up the third group, looking at whether the data treats different groups of people evenly. Privacy rounds out the list, checking whether people in the data can be identified and whether sensitive details are protected.

“It is essential that data quality is measurable, especially for organisations who need to establish whether its data is fit to essential intents, like it would be the case of trustworthy AI,” Diego Lopez, Chair of the ETSI Technical Committee DATA, said.

“ETSI’s standardised metrics provide a common language for assessing data quality, giving quantitative evidence as to whether a dataset is fit for its intended purpose. This lays important groundwork for more consistent and repeatable approaches to data quality assessment, as AI and data-driven technologies continue to evolve,” Lopez added.

To test the metrics, the researchers behind TR 104 180 applied them to two public datasets. One held sensor readings from aircraft engines. The other was a US census dataset, long used in machine learning research, built to predict whether someone earns over 50,000 dollars a year.

The engine data held up well. It was complete, accurate, and steady over time. The census data raised more concerns.

A gender gap in the numbers

The researchers looked at the census data for bias between men and women. About 31 percent of men in the dataset were marked as high earners, compared with about 11 percent of women. That gap is close to three times over, which TR 104 180 treats as a warning sign for bias.

Two privacy problems in one dataset

The census data also failed on privacy, in two separate ways. First, looking at just four details together, age, race, sex, and country, was enough to single out specific people in the dataset. Some individuals could be identified on their own, which the report flags as a serious risk.

Second, when the researchers checked whether sensitive fields were protected, they found personal information stored in plain text, with no masking or encryption in place.

TR 104 180 treats both findings as data quality failures. Anonymity and confidentiality appear on the same list as accuracy and completeness, scored with the same kind of formulas.

TR 104 180 was developed with a working group that included Sejong University, EGM, TTA, Daejeon University, and CNIT. The group also built an open-source tool that scores any dataset against the 18 metrics.


from Help Net Security https://ift.tt/qZiyHg5