Three independent security researchers say they used Anthropic's Claude to break into internal OpenAI systems in less than 72 hours, gaining access to employee accounts and the company's closely guarded GitHub repository — then walked away with a bug bounty payment rather than stolen data.
The findings, first reported by the Wall Street Journal and subsequently covered by The Verge, The Guardian, CBS News and a wave of international outlets, describe one of the most striking demonstrations yet of how quickly frontier AI models can accelerate offensive security work. The team, operating under the name Hacktron, used Anthropic's Claude Opus 4.8 and Claude 5 to map and penetrate OpenAI's defenses.
How the breach unfolded
According to the accounts, the researchers did not attack OpenAI's core model infrastructure head-on. Instead, they worked through the periphery. The entry point was Discourse, the third-party platform that hosts OpenAI's community forum — a service that sits outside the company's most hardened perimeter but still connects to its identity and access systems.
From there, the team was able to compromise OpenAI employee accounts. That foothold gave them access to the company's GitHub environment, including a repository known internally as "Monorepo." The Wall Street Journal's sources described the repository as containing what amounts to OpenAI's algorithmic secrets — the accumulated engineering work behind its models.
"OpenAI's algorithmic secrets" — the Wall Street Journal's characterization of what the Monorepo repository is said to contain.
Crucially, the researchers say they stopped short of exfiltrating internal code. To prove they had genuinely obtained access, they instead pushed a pull request from an employee's Codex account — a verifiable breadcrumb that could be traced back to the compromised identity without removing anything of value.
A bounty, not a crime
The operation was framed as ethical hacking. The Guardian described it as an "ethically hacked" exercise, and the researchers reported their findings rather than exploiting them. According to Indian media coverage, the team — which includes Indian-origin researchers — was paid roughly Rs 6.2 lakh, or about $7,000, for their work.
That payment places the exercise squarely inside the bug bounty economy, where companies incentivize outsiders to find flaws before malicious actors do. OpenAI operates one of the more prominent bounty programs in the industry, and the payout suggests the company treated the findings as a legitimate vulnerability disclosure rather than a hostile intrusion.
Even so, the details are uncomfortable. The researchers relied on a third-party vendor — Discourse — to reach a company whose entire brand rests on being at the frontier of technology. Supply-chain and periphery services remain among the most reliable ways into even well-defended organizations, and automated reasoning tools make finding those seams dramatically faster.
The strategic irony
Perhaps the most notable wrinkle is who built the tool. Claude is a product of Anthropic, OpenAI's most direct commercial rival in the race to build frontier models. In this episode, one competitor's AI was used to probe the other's defenses — a vivid illustration of how AI capability diffuses outward, often faster than the companies that create it can anticipate.
The coverage across outlets reflects differing emphases. The Verge and The Information framed the story around the technical mechanics and the target: the breach of a leading lab's code repository. The Guardian led with the ethical-hacking framing. Indian outlets, including Financial Express, emphasized the speed — 72 hours — and the bounty earned by researchers of Indian origin. Aggregators such as MSN syndicated the wire reports with headlines oscillating between "hacked into OpenAI" and "used Claude to hack ChatGPT," a looseness that blurred the distinction between the company and its flagship product.
Anthropic's own security problem
The story lands amid broader anxiety about how AI systems can be misused. Separately, Anthropic has said it disrupted efforts involving bioweapons research, Russian hacking operations and misuse of Claude by actors in China. The pattern is consistent: the same models that help defenders find flaws also lower the barrier for attackers, and the line between the two depends almost entirely on intent.
What it means
- For AI labs: The attack surface extends well beyond the model itself, into forums, identity providers and third-party SaaS tools.
- For defenders: Autonomous agents compress the timeline of reconnaissance and exploitation from weeks to days.
- For the industry: Cross-vendor AI use — one lab's model testing another's infrastructure — is now a normal feature of security research.
Whether the episode becomes a cautionary tale or a validation of bug bounty culture may depend on what OpenAI changes next. Neither OpenAI nor Anthropic has publicly detailed the full scope of the incident, and the researchers' account rests largely on reporting by the Wall Street Journal and its sources. What is clear is the asymmetry the episode exposes: 72 hours, three people, and a competitor's chatbot were enough to reach the doorstep of one of the most consequential codebases in the world.




