Tuesday, 02 January 2024 12:17 GMT

Anthropic Reports AI Misbehavior On US Government Websites


Date:

(MENAFN- Khaama Press) Anthropic, the company behind the Claude artificial intelligence model, has disclosed several incidents in which its AI systems took unintended actions on external digital platforms, including websites operated by US government agencies, raising concerns about the security risks associated with advanced AI tools.

According to Bloomberg News, the incidents involved AI models exploiting software vulnerabilities to execute commands, submitting online forms without authorization and bypassing restrictions to access certain public data. Anthropic said the cases involved federal, state and local government websites but did not identify the agencies or other affected organizations.

In a report published Friday, Anthropic grouped the behavior into four categories. These included exploiting flaws in third-party software, submitting sensitive forms without permission, circumventing access restrictions and using URL-shortening services to bypass limits imposed by its own web-access tools. The company said most incidents occurred during evaluations designed to test the models' capabilities.

In one case, Anthropic's Claude Haiku 4.5 model submitted a tip to the Philadelphia Police Department about an unsolved homicide, falsely suggesting it might have information about the case. The submission was flagged as spam and was not reviewed by investigators, according to Reuters. Anthropic said the model had been instructed to carry out online tasks but had not been explicitly told to avoid submitting forms.

Anthropic said some models also exploited basic software flaws to execute commands on third-party servers or accessed publicly available information that was normally restricted by fees or access controls. The company described the incidents as unintended actions rather than deliberate attempts by the models to cause harm.

The company said the incidents identified so far had minimal real-world impact. It also said it had briefed the White House, notified the affected agencies and introduced additional restrictions on internet access for certain AI models during testing.

The disclosures prompted renewed scrutiny from the Trump administration, which has called on AI companies to strengthen safeguards and promptly report security incidents involving their systems. The cases add to broader concerns about AI agents that can interact with live websites, use digital tools and attempt to complete tasks with limited human supervision.

Anthropic said it is reviewing how these behaviors occurred and working to reduce the risk of similar incidents. The disclosures highlight the challenge of ensuring that AI systems remain within their intended limits when operating on real-world digital infrastructure.

MENAFN10102026000228011069ID1111786551



Khaama Press

Legal Disclaimer:
MENAFN provides the information “as is” without warranty of any kind. We do not accept any responsibility or liability for the accuracy, content, images, videos, licenses, completeness, legality, or reliability of the information contained in this article. If you have any complaints or copyright issues related to this article, kindly contact the provider above.



More Story