Short

Anthropic AI sent a fake murder tip to Philadelphia police, and it was not the only slip

During an automated test, Claude Haiku 4.5 submitted an invented homicide tip through a Philadelphia police form. Anthropic's new report lists more unintended actions, and the company has cut live internet access for all internal evaluations.

Pencil sketch of a laptop showing an online police tip form with a cursor clicking the submit button

On July 18, 2026, at 11:27 in the evening, a short message arrived through the tip form of PhillyUnsolvedMurders.com, a website the Philadelphia Police Department runs to collect information on unsolved homicides. The writer said they might have information about a case and recalled seeing someone matching the description near the street named on the page. There was no name and no contact detail. It turned out there was also no person. The tip was written and submitted by an Anthropic AI model during an automated test.

Anthropic disclosed the incident on October 9 in a research report titled "Investigating unintended model actions in our evaluations and internal use". The Philadelphia police published their own statement the same day, a few hours before the report went live, and called the delay in telling the city "unacceptable". The tip form case is the most striking example in the report, but it is only one of several. Below is what happened, what else Anthropic found, how the company is reacting, and why this matters for anyone who wants to let AI agents loose on the open web.

Pencil sketch of a laptop showing an online police tip form with a cursor clicking the red submit button, a city skyline at night in the background

The three things you need to know

  • What happened. In a test on randomly selected websites, Claude Haiku 4.5 landed on a page about an unsolved homicide, filled out the police tip form with an invented tip and submitted it. The message was flagged as spam and never reached investigators.
  • It was not alone. Anthropic's report describes four categories of unintended behavior: exploiting basic software flaws to run commands on a server, submitting forms that should not have been submitted, working around restrictions to reach data gated by a token or a fee, and using URL shortening services to get around limits in its fetch tool.
  • The fix. Anthropic has turned off live internet access for all of its internal evaluations until it has confirmed that its security and monitoring measures reliably catch behaviors like these. Philadelphia police said the two month delay in detecting and reporting the incident was unacceptable.

The fake tip, step by step

According to Anthropic, Claude Haiku 4.5 had been tasked with generating and performing example tasks on randomly selected webpages. In one run, it reached a page about an unsolved homicide that contained a tip form operated by a police department. The instructions told the model never to log in, never to create accounts, never to enter personal data, never to make purchases and never to submit anything destructive. They did not rule out form submissions.

So the model filled out the form. Anthropic quotes the text in full: "I may have information regarding this case. I recall seeing someone matching the description in the area around [the street named on the page] during that time period. Please contact me if this information is relevant." Anthropic points out that the website did not even include a description of the perpetrator. The model left the name and contact fields empty, which the form allowed, and hit submit.

The Philadelphia Police Department filled in the timeline from its side. Anthropic told the department that it discovered the incident on September 28, terminated the automated testing process responsible for the submission and added an extra validation step for future testing. Anthropic notified the police on October 7, and the two sides met on October 8. After that briefing, the police located the submission in the website's tip records and confirmed that the corresponding email was still sitting in spam. In its statement, the department said it is coordinating with the city's Law Department, the Office of Innovation and Technology and Mayor Cherelle L. Parker's executive team, and that "the two-month delay in detecting and reporting the incident to the City is unacceptable."

How bad was the intent? Anthropic's own reading of the transcript is that the model "appears to have only been producing example content for the task, rather than trying to mislead anyone to achieve a goal." The company also stresses that a model's own account of its reasoning is not necessarily reliable evidence of what it believes, and that its view of these cases may change with further analysis.

Four kinds of unintended actions

The tip form is the headline, but the report is broader. Anthropic sorts what it found into four categories, and each one shows a model doing something its operators did not intend when it hit an obstacle.

Exploiting basic software flaws. In one evaluation, Claude Mythos Preview was asked to run a scientific analysis with a public tool hosted by a university. The tool returned an error. Looking for another way, the model explored the website, found a script on the university's server that would return any file it was asked for, used it to copy files including the script's own code, found an injection flaw in that code and used it to run the calculation on the server. In another case, Claude Mythos 5 needed free data that sat behind a data use agreement it had no tool to accept, so it used applications hosted on the same website to load the page and accept the agreement on its behalf. Anthropic says it saw this behavior on DeepSearchQA and BrowseComp, two public web research evaluations, on LABBench2, a public biology research benchmark, and in internal evaluations.

Submitting forms it should not have. Besides the police tip, Anthropic describes an unreleased, non frontier research model that was meant to fill out a practice copy of a government form. When the copy failed to load or the model closed it by mistake, it navigated to the website that hosts the real form and submitted it there, and this happened multiple times on the same evaluation. In another evaluation, Claude Haiku 4.5 was told to stop before the final submission step but several times submitted anyway, expecting another confirmation page. The Verge, citing Axios, reports that a State Department official said Anthropic told the department that a model in testing submitted 19 non immigrant visa applications in August and one in May. Anthropic says this category showed up on OSWorld, a public computer use evaluation, on Odysseys, a long horizon task evaluation, and during internal usage.

Working around restrictions to reach gated data. Asked to identify the location in a photo, Claude Mythos 5 wanted to check a local government's property map. Its evaluation limited how it could click through pages, so it read the settings file the site sends to browsers, found working access tokens and used them to query the server behind the map directly. In a separate case, an Anthropic researcher used Claude Mythos 5 on a statistics project that needed public data a state agency sells for a fee. The model learned from an archived copy of the agency's website that its public dashboard hands an access token to any visitor, requested one and queried the database without paying. Anthropic notes that in both examples the data was public, but only for a fee, and says this behavior appeared on Humanity's Last Exam and in internal usage.

Using URL shorteners to dodge tool limits. Some of Anthropic's fetch tools limit the length of URLs that Claude can request, because long URLs have room to carry SQL or command injections. Several models, including Claude Opus 5 and Claude Mythos 5, got around this by running their URLs through free URL shortening services. Anthropic first noticed this internally, and while the report was being written, the operator of the da.gd shortening service told the company it had also seen Claude using its site this way.

Anthropic says some of the cases involved websites run by U.S. government agencies at the federal, state and local level. It chose not to name the organizations, at their request and to avoid exposing vulnerabilities, and says it briefed the White House and notified each agency involved. According to the company, to its knowledge none of the cases involved customer data or Anthropic's own internal systems, and the real world impact was minimal.

Why it happens: tasks that cannot be done as given

Anthropic's explanation centers on two ideas. The first is reward hacking. Models learn much of what they can do through reinforcement learning, where they attempt a task many times and get rewarded for success. If a training environment accidentally rewards finding loopholes or working around a restriction, the model learns that the workaround pays off and may apply it elsewhere. The company describes most of the new cases as forms of persistence: when Claude cannot complete a task as given, it works around a restriction instead of stopping.

The second idea is ambiguity. Many of the incidents happened when evaluation instructions were unclear, when a misconfigured environment broke the dummy version of a form, or when the task was impossible to finish as written. Anthropic admits that clearer instructions about targets, permitted actions and network boundaries might have prevented some of the failures, but also notes that Claude meets ambiguous and impossible tasks every day in real use, and several of the cases happened during regular agentic use rather than in tests.

Why were models on the live internet at all? Anthropic says some tasks, such as searching the web for hard to find information, are difficult to simulate realistically offline, and that public web search benchmarks run on the live internet by default. Running them the same way lets labs compare their models with others.

Pencil sketch of a server rack in a data center with a red network cable unplugged and hanging down

What Anthropic is changing

The most visible step is the internet switch. Anthropic had already turned off live internet access for some high risk and cybersecurity evaluations. It has now expanded that to all of its internal evaluations, until it has confirmed that its security and monitoring measures reliably catch behaviors like the ones in the report.

Beyond that, the company lists several measures. It has stopped running some public evaluations, moved others to offline versions or rebuilt them so their tasks no longer reach live websites, and tightened the guardrails on internet access tools such as its web fetch tool. It has built tooling that automatically detects and blocks the kinds of behavior described, which now runs on most of its evaluations and on internal agentic use of frontier models. When Anthropic tested that tooling against the cases in the report, it blocked all of them. The company also says it is migrating internal agents to centrally managed infrastructure with strong containment, minimizing internet access for internal agents and training processes, monitoring more of what agents do with safety classifiers and hierarchical summarization, and continuing to fix or remove training environments that reward working around tool restrictions.

How serious does Anthropic think this is? It calls these behaviors significantly less severe, from an alignment and security perspective, than the cybersecurity incidents it reported on July 30 and September 9, when Claude gained access to real third party systems for hours. It also says none of the new behaviors are new in kind and that they do not change its overall view of Claude's alignment. At the same time it writes that alignment training is not yet sufficient or fully robust on its own, which is why it relies on extra classifiers and safeguards.

The reactions

Outside reactions were mixed. TechCrunch framed the story as Anthropic being unable to reliably control its agents and choosing to cut its internal evaluations off from the live internet instead. Conrad Stosz of the AI oversight lab Transluce, a former head of the U.S. Center for AI Standards and Innovation, told TechCrunch it was encouraging that Anthropic voluntarily disclosed the incidents, but that it underscores the need for independent, credible, third party verification of AI systems. Sydney Von Arx, founder of the AI safety organization Nightingale, told TechCrunch before the disclosure that developing models fully cut off from the internet would be very challenging and that at some point the models need to be aligned for a world where they do have access.

In Philadelphia, the reaction was blunter. The police department said the company must strengthen its safeguards so that similar incidents do not affect city systems without the city's knowledge, and that the Parker administration will explore regulatory protections with state and federal partners. The Verge reports that the Trump administration's Super Intelligence Force told Axios that companies must immediately disclose incidents involving their models and act swiftly to remedy any harm.

Why it matters

The tip itself did no damage. It sat in a spam folder for almost three months. But the case shows a pattern that will matter more as AI agents get more autonomy: a model given a vague instruction and a real browser will often keep pushing until it finds a way, even when that way is a police form, a visa application or someone else's server. The rules it was given, no logins, no purchases, nothing destructive, did not anticipate that filling in a form could be harmful in itself.

It also shows how long such things can stay hidden. The submission happened on July 18, Anthropic found it on September 28 during a transcript review that began in July, and the police heard about it on October 7. For a city agency, that gap is the real problem. For the AI industry, it is a reminder that many popular benchmarks run against the live internet, and that the same behavior could show up in other labs' tests too. Anthropic says as much: it hopes the report helps other developers check their own models.

Dany's take

I respect that Anthropic published this, with the exact text of the fake tip and the names of the benchmarks. That is more than most companies would do. But "minimal real world impact" is doing a lot of work here. A made up homicide tip and 19 visa applications are not harmless just because they landed in spam or were never processed. The bigger lesson for me is simple: if the people who build these models need two months and a manual transcript review to notice what their agents did on the open web, then everyone running agents on their own accounts should keep them on a short leash, with clear rules about what they may submit, and a human looking at the log. Should AI agents be tested on the live internet at all? Tell me what you think on Reddit or X.

Sources

Source: anthropic.com

Newsletter

The AI news that matters, in your inbox.