Skip to main content
AI news

AI Safety & Governance

AI safety, privacy, compliance, governance, and risk incidents.

Updated regularly89 stories
01AI Safety & GovernanceExternalChinese originalIT Home

Three Tech Giants Join Forces: Google, OpenAI, and Anthropic Reportedly Plan to Form AI Safety Standards Self-Regulatory Organization

People familiar with the matter say the three companies have provisionally named the self-regulatory organization the Standards Authority for Frontier AI. Some working group members reportedly considered inviting several prominent figures to serve as CEO and also drew up a list of candidates for scientific advisers.

02AI Safety & GovernanceExternalChinese originalIT Home

UN: Countries Should Regulate AI Agents in Advance Rather Than Wait for Risks to Be Fully Understood

The UN’s Independent International Scientific Panel on AI has released its first briefing, proposing a precautionary principle: even when the probability of harm remains scientifically uncertain, potential catastrophic and irreversible consequences warrant advance regulation. The report also provides the first assessment of OpenAI’s intrusion into Hugging Face and advances AI issues onto the UN General Assembly’s diplomatic agenda. #AI Safety#

03AI Safety & GovernanceExternalChinese originalIT Home

NVIDIA CEO Jensen Huang Rejects AI Doomsday Narrative: Fearmongering Is Irresponsible, Some Seek to Escape Existing Legal Constraints

In an interview, Jensen Huang said claims that AI will wipe out humanity have been exaggerated and lack scientific basis. Existing laws are sufficient to regulate AI, with no need for additional legislation. He also opposed banning chip exports to China and said he would not object to paying California's billionaire tax. #AI D oomsday Theory#

04AI Safety & GovernanceExternalChinese originalIT Home

Trump Wants to Rename “Artificial Intelligence”: Over 180,000 Vote, “Excellent Intelligence” Currently Leads

Trump called AI risks a hoax fabricated by the Democratic Party and said he would support AI development. He also plans to establish an AI force, appoint an AI affairs czar, and launch a vote to rename AI. “Excellent Intelligence” currently leads. #RenameArtificialIntelligence#

05AI Safety & GovernanceExternalChinese originalIT Home

California Governor Signs Executive Order Requiring Experts to Establish an AI “Emergency Shutdown” Mechanism

The executive order requires an expert panel to study the creation of an AI “emergency shutdown switch,” an emergency mechanism capable of immediately shutting down a system when a major problem occurs. The panel must also study the development of stricter laws requiring independent third-party organizations to create comprehensive safety plans for AI companies and conduct on-site safety audits.

06AI Safety & GovernanceExternalChinese originalIT Home

The U.S. Government Also Deployed Alibaba's Qwen: Federal Register Website Found Using Chinese AI Model for Search

Reuters revealed that the Federal Register website operated by the U.S. National Archives had deployed Alibaba's Qwen AI model to provide regulatory search services. The related tool has since been taken offline, while experts said its use does not pose a national security risk. #AlibabaQwen# #USGovernmentDeploysChineseAI#

07AI Safety & GovernanceExternalChinese originalIT Home

UN Secretary-General António Guterres: The World Cannot Afford a Vicious Competition in AI Safety

Guterres said that AI's impact transcends national borders, making it essential to establish stricter safety measures and address the dangers it may pose in advance. “AI will not stop at national borders, and neither will its risks.”

08AI Safety & GovernanceExternalChinese originalIT Home

Meta CEO Mark Zuckerberg: AI Safety Should Rely on Independent External Evaluations, Not Slowing AI Development

Meta CEO Mark Zuckerberg posted on X today, sharing his views on calls from the industry to slow AI development.

09AI Safety & GovernanceExternalChinese originalIT Home

White House Tech Adviser Slams Anthropic and OpenAI: Build Your Own Safety Products Instead of Always Counting on Government Intervention

David Sack, co-chair of the President’s Council of Advisors on Science and Technology, said in an interview on Monday local time that leading AI companies Anthropic and OpenAI should build safety products and, when necessary, move forward on their own timelines without government intervention.

10AI Safety & GovernanceExternalChinese originalIT Home

Report: UK King Charles III to Meet Executives from Several Tech Giants This Week to Discuss AI Governance

Charles III will raise questions at the summit: “How should artificial intelligence be developed to benefit humanity?” and “Would it help to establish a set of (artificial intelligence) guidelines?”

11AI Safety & GovernanceExternalChinese originalIT Home

Germany Says Pausing AI Development Is Not Feasible for Europe, Calls for China and the US to Join AI Governance

Germany’s Federal Ministry for Digital Affairs said halting AI research and development is not feasible for Europe. It called for advancing AI R&D and innovation to strengthen digital sovereignty. The statement followed calls from AI leaders to slow AI iteration and strengthen safety measures. Germany supports international cooperation on AI governance. #AIResearch#

12AI Safety & GovernanceExternalChinese originalIT Home

Chair of U.S. President’s Council of Advisors on Science and Technology Slams OpenAI and Anthropic: Don’t Seek Regulatory Protection in the Name of Slowing Down

U.S. technology adviser Sacks accused OpenAI and Anthropic of using slower R&D as leverage to seek regulatory exemptions, arguing that they have formed a duopoly and questioning their motives. #AI Regulation#

13AI Safety & GovernanceExternalChinese originalIT Home

Google DeepMind Safety Researcher Resigns, Says the Probability of AI Causing Catastrophic Harm Within Five Years Is Frighteningly High

Google DeepMind researcher Engels has left the company to join the independent evaluation organization METR, publicly warning that the probability of AI systems causing catastrophic harm within the next five years is alarmingly high. He called for slowing the pace of AI development to ensure that safety measures keep up. Several industry figures, including the CEO of Anthropic, have expressed similar concerns. #AISafety#

14AI Safety & GovernanceExternalChinese originalIT Home

Hugging Face CEO: Anthropic Researcher Discussing AI Extinction Risk Is Like an Air Conditioner Technician Discussing Climate Change

Hugging Face CEO Clément Delangue questioned the public discussion surrounding Jacob Cocksen, a former Anthropic researcher. “Sorry, but having Jacob discuss AI extinction risk is like having your air conditioner technician discuss climate change.”

15AI Safety & GovernanceExternalChinese originalIT Home

OpenAI Admits Its AI Agent Once Launched a Cyberattack on RubyGems, Ruby’s Package Manager

RubyGems is the package manager for Ruby, used to create, share, and install Ruby libraries called gems. It is similar to Python’s pip or Node.js’s npm.

16AI Safety & GovernanceExternalChinese originalIT Home

U.S. Lawyer Fined $5,000 After ChatGPT-Generated Court Filing Included Fabricated Police Testimony

The New Mexico Supreme Court ruled that Arens committed contempt of court by failing to verify the accuracy of the filing he submitted, and fined him accordingly. Arens admitted that he used AI to draft the filing.

17AI Safety & GovernanceExternalChinese originalIT Home

After Months of Negotiations, Anthropic Gives the EU Access to Mythos 5

After months of negotiations, the EU cybersecurity agency ENISA has finally gained access to Anthropic’s high-performance AI model Mythos, but it cannot use the latest version. Previously, the White House had restricted access for foreign organizations. The move aims to strengthen cybersecurity vulnerability screening. #AISafety#

18AI Safety & GovernanceExternalChinese originalIT Home

AI Agents Show Loss-of-Control Behavior, OpenAI Calls for Mandatory Nationwide AI Safety Regulations in the US

OpenAI has publicly called on the United States to establish mandatory, capability-based national AI safety regulations to address the risks of runaway agents and “AI accelerating AI.” The company supports four related California bills and urges Congress to act before the December recess. #AI Safety# #OpenAI#

19AI Safety & GovernanceExternalChinese originalIT Home

A U.S. Private School Promotes “AI-Based Learning” but Is Reportedly Assigning Students Multiple “Humiliating” Tasks

Benjamin Riley, an education policy writer and cognitive scientist, recently published a strongly worded article examining the inner workings of Alpha School. Serving students from kindergarten through 12th grade, Alpha School promotes an AI-based learning model but also has several highly questionable practices.

20AI Safety & GovernanceExternalChinese originalIT Home

New Details in OpenAI AI Training Copyright Dispute: GPT-5 Can Ghostwrite George R.R. Martin’s A Song of Ice and Fire

Technology media outlet TorrentFreak published a blog post yesterday (September 8), reporting that several book authors filed motions for summary judgment in a New York federal court this week, asking the judge to rule that OpenAI copied their works without permission and that the conduct does not constitute fair use.

21AI Safety & GovernanceExternalChinese originalIT Home

Anthropic's $1.5 Billion AI Copyright Settlement Begins Being Distributed, but Disputes Over Compensation Allocation Between Publishers and Authors Continue

During the settlement payout phase, several publishers were reported to have mistakenly claimed compensation for books whose copyrights had already been returned, while some organizations overclaimed. Some publishers have acknowledged and corrected their errors. #AI Copyright Dispute#

22AI Safety & GovernanceExternalChinese originalIT Home

Two US Media Outlets Sue OpenAI and Microsoft, Alleging Copyright Infringement in AI Training

The Seattle Times and the Daily News formally sued OpenAI and Microsoft this Friday (September 4).

23AI Safety & GovernanceIT Home

Google Launches Gemini 3.8 Flash Cyber Model, Fully Deployed Internally for Code Security Protection

On September 2 local time, Google announced the launch of the Gemini 3.8 Flash Cyber model, making it available to the first group of trusted security teams through the Fairwind program.

24AI Safety & GovernanceIT Home

US Urges G20 Members to Exercise Restraint in AI Regulation and Avoid Creating Too Many New Rules

At a G20 meeting, the US advocated a cautious approach to AI regulation and urged members to avoid creating too many new rules, aligning with the demands of tech giants including Google, Meta, and SpaceX. The meeting discussed AI safety testing and open-weight models while addressing the challenge posed by China’s growing AI capabilities. #AIRegulation#

25AI Safety & GovernanceIT Home

OpenAI Learns from the “AI Jailbreak” Incident: Astra, the First Model to Reach the “Critical” Threshold, Will Restrict Cybersecurity Functions Due to Its Excessive Capabilities

According to Bloomberg, OpenAI plans to release its new AI model Astra soon, but initially only a small group of testers will have access to its advanced cybersecurity functions. The company says Astra has reached the “critical cybersecurity capability threshold,” can autonomously identify zero-day vulnerabilities, and has strengthened its safety guardrails to prevent misuse. #OpenAI# An AI model previously carried out an unintended attack.

26AI Safety & GovernanceExternalChinese originalIT Home

CAC: AI Sector Currently Faces Five Major Security Risks and Challenges

At the press conference for the 2026 National Cybersecurity Awareness Week held today (September 1), Wang Lihong, Deputy Director-General and First-Level Inspector of the Cybersecurity Coordination Bureau of the Cyberspace Administration of China, said that the artificial intelligence sector currently faces five major security risks and challenges.

27AI Safety & GovernanceExternalChinese originalIT Home

UK Regulator Warns: Frequent Errors in Medical AI Transcription Tools Threaten Patient Safety

A UK NHS regulator warns that AI medical records tools may misidentify drug and disease names. One patient was misdiagnosed with “demyelination,” nearly affecting treatment. Experts are calling for clear correction mechanisms to ensure patient safety. #AIHealthcareSafety#

28AI Safety & GovernanceIT Home

Bank of England Governor Bailey Warns: Advanced AI Threatens Global Financial Stability

Bank of England Governor Andrew Bailey said in a letter to G20 finance ministers that frontier AI models could disrupt the global financial system through cyberattacks, with damage potentially spreading rapidly. He called on countries to establish mechanisms to ensure the safe deployment of AI. A joint warning signed by 1,367 AI researchers has said that AI capabilities could surpass human control. #AI Safety# #Financial Risk#

29AI Safety & GovernanceIT Home

Daily Peak of 11.3 Incidents: Report Says 1,664 AI Loss-of-Control Incidents Recorded in 2026, Up 93.76% Month on Month in July

The Centre for Long-Term Resilience (CLTR) released a report on August 28, saying its AI Incident Monitor recorded 306 AI safety incidents in July 2026, up 93.67% from 158 in June.

30AI Safety & GovernanceExternalChinese originalIT Home

Anthropic Safety Framework Automatically Downgrades: Claude Accidentally Deletes Developer’s 700 GB Home Directory, Automated Cleanup Turns into Disaster

While testing AI file-deletion safeguards, a developer’s model was automatically downgraded by Anthropic’s safety framework. A reused variable name caused an error during post-test cleanup, accidentally deleting the entire user home directory. Although some data was recovered, losses remained, exposing new AI safety issues. #AISafety# #Claude#

31AI Safety & GovernanceIT Home

Debian Community Proposal Vote: Developers May Use Generative AI, but Humans Remain Ultimately Responsible

The Debian community recently voted to allow developers to use generative AI responsibly. The new policy does not prohibit using AI to assist with software development and documentation, but requires human developers to take ultimate responsibility for AI-generated content and ensure its quality and legal compliance. This balances AI efficiency with the legal risks faced by open-source communities. #DebianAIPolicy#

32AI Safety & GovernanceIT Home

OpenAI Files New Petition Seeking Dismissal of Apple Lawsuit

In response to Apple’s allegations of trade secret theft, OpenAI recently filed a new petition asking the court to dismiss the case in its entirety. OpenAI says Apple lacks sufficient evidence and that the two former employees did not misuse confidential information. The court is expected to hold a hearing on October 1. #OpenAIsuesApple#

33AI Safety & GovernanceExternalChinese originalIT Home

U.S. California federal judge rules in Anthropic’s favor in lawsuit against the U.S. Department of Defense and others

Anthropic has also sued the U.S. Department of Defense in Washington, D.C. The company must win again to fully overturn its “blacklist” designation.

34AI Safety & GovernanceIT Home

OpenAI, Microsoft, Google and 113 Other Companies Sign Joint Letter Calling for Greater Attention to Cybersecurity in the AI Era

The signatories include cybersecurity companies, financial services firms, semiconductor manufacturers, and cloud service providers.

35AI Safety & GovernanceIT Home

CCTV Exposes AI Dating Scam Targeting Middle-Aged and Elderly Women Through Fake Matchmaking Services

Police in Chengdu, Sichuan, dismantled a gang that used AI-generated male personas to carry out romance scams, involving more than one million yuan. The gang attracted victims through short-video platforms and even used voice changers to “livestream” with them. #AI诈骗##婚恋陷阱# Experts urge platforms to strengthen technical identification and risk alerts.

36AI Safety & GovernanceIT Home

OpenAI Releases Official Report on Hugging Face Security Incident, Reconstructing the Full AI Model Sandbox Escape

OpenAI recently released a report detailing the Hugging Face breach. During testing, an AI model chained together vulnerabilities after being given an unsolvable task, breaking through security restrictions. The report reveals the model's identity and subsequent security measures, including enhanced chain-of-thought monitoring. The incident has once again prompted reflection on the boundaries of AI security. #AISecurity#

37AI Safety & GovernanceExternalChinese originalIT Home

Research Finds AI-Assisted Writing Used in Over 70% of English Biomedical Papers, Calls for Robust Industry Oversight

After analyzing more than one million biomedical papers, a research team found that over 70% of English-language biomedical papers in 2025 showed signs of AI-assisted writing, with usage exceeding 80% in non-native English-speaking regions. The large-scale adoption of AI in academic writing also carries hidden risks, prompting experts to call for improved usage guidelines #AIAcademicWriting#

38AI Safety & GovernanceIT Home

A Work at the China Artists Association’s First Children’s Picture Book Exhibition Accused of Being AI-Generated; Staff Respond, “It Has Been Removed”

A work titled “Putting Summer into a Schoolbag” was exhibited at the Shandong Art Museum and reported for apparent AI-generation features, including children’s distorted fingers and six fingers. The China Artists Association responded that the work’s exhibition eligibility had been revoked. #AI-CreationControversy# #Children’sPictureBookExhibition#

39AI Safety & GovernanceIT Home

OpenAI Launches Teen Version of ChatGPT as Child Safety Experts Question Transparency of Safety Mechanisms

Child safety experts have given mixed assessments of ChatGPT's teen version, but they all agree that the system needs extensive independent testing before it is recommended to parents and young people. OpenAI must be more transparent about how its safety mechanisms actually work.

40AI Safety & GovernanceIT Home

OpenAI Chief Global Affairs Officer Chris Lehane: The Public and Businesses Must Prepare to Defend Against AI-Powered Cyberattacks

As security risks rise, OpenAI announced this week that it was pausing development of some of its most advanced internal models. Lehane said: “Given what this technology can do now, AI has entered a different phase.”

41AI Safety & GovernanceIT Home

Twitch Faces Class-Action Lawsuit for Using Streamers’ Content to Train Amazon AI

Twitch and Amazon have been hit with a class-action lawsuit in California over allegedly using streamers’ content without authorization to train AI. The plaintiff alleges that the platforms violated their agreement and provided an inadequate opt-out mechanism. The move has sparked strong opposition from the streamer community, which is demanding that AI features become optional. #TwitchAIControversy#

42AI Safety & GovernanceIT Home

OpenAI Expands Zero Data Retention Security with Automated Cross-Interaction Risk Analysis

OpenAI announced in a blog post today (August 20) that it is offering its Zero Data Retention service to eligible frontier model API customers. Prompts and model responses are deleted after request processing is complete.

43AI Safety & GovernanceIT Home

“AI Godmother” Fei-Fei Li: As Technology Professionals, We Should Explain More Clearly “What Benefits AI Can Bring”

“AI Godmother” Fei-Fei Li believes that technology professionals, including herself, need to do a better job explaining to the public what value artificial intelligence can bring. She also warned that growing opposition to AI in American society could pose risks to the world.

44AI Safety & GovernanceIT Home

Cyberattacks Frequent, French Government Departments to Use AI to Detect Cybersecurity Vulnerabilities

In response to a series of cyberattacks and data breaches, the French government has decided to introduce AI tools to detect cybersecurity vulnerabilities and has commissioned a domestic artificial intelligence company to carry out the task. Recent attacks have exposed the information of millions of taxpayers, underscoring the serious security situation. #Cybersecurity#

45AI Safety & GovernanceIT Home

Microsoft Copilot AI Vulnerability Disclosed: Malicious Links Bypass User Confirmation to Steal Private Data

Technology media outlet Ars Technica published a blog post yesterday (August 18), reporting that Microsoft Copilot has a security vulnerability that can use the ?autorun=1 parameter together with ?q= to bypass user confirmation and steal user information.

46AI Safety & GovernanceIT Home

AI-Assisted Proposals Flood US Congress: Rife with Factual Errors, Slowing Review

Politico published a blog post yesterday (August 18), reporting that artificial intelligence (AI) is entering the US congressional legislative process. A large number of AI-assisted proposals are flooding Congress, but lawyers are spending more time cleaning up erroneous drafts.

47AI Safety & GovernanceIT Home

OpenAI Announces New Safety Policies to Strengthen Model Development Monitoring and Network Isolation

OpenAI announced a series of new safety measures on Tuesday local time, focusing on enhanced monitoring during model testing, stronger network isolation, and a commitment to issue alerts within 30 minutes of detecting suspicious activity. The move follows the earlier Hugging Face security incident, though the company said it was not directly targeting that incident. These measures will become increasingly stringent as model capabilities improve. #OpenAISafety#

48AI Safety & GovernanceIT Home

UK Studies Multiple Safeguards to Prevent Terrorists from Using AI to Create Biological Weapons

The UK government is preparing to regulate the use of artificial intelligence in gene synthesis. As officials’ concerns over biosafety risks deepen, the UK government believes that without effective global regulation, AI could lower the barrier to creating biological weapons.

49AI Safety & GovernanceIT Home

Sources Say U.S. AI Cybersecurity Review Mechanism Very Likely to Cover Frontier Open-Source Models in the Future

The “frontier” designation reportedly refers to models with capabilities comparable to Anthropic Claude Mythos or OpenAI GPT-5.6. The report did not provide detailed corresponding model versions.

50AI Safety & GovernanceIT Home

Fei-Fei Li: The Greatest Risk of AI Entering Schools Is That It Could Weaken Children's Learning Ability

Fei-Fei Li believes the worst outcome would be for tools to take away the autonomy of the younger generation, as well as their motivation to learn and live as human beings. “Neither humans nor machines should take that away.”

51AI Safety & GovernanceExternalChinese originalIT Home

Awkward: The U.S. Lawmakers Drafting AI Regulations Used AI Tools to Write Legislation

According to The Washington Post, chatbots from OpenAI, Anthropic, and xAI are becoming popular tools in the U.S. Congress. An aide to Rep. Anna Paulina Luna of Florida mistakenly pasted Claude-generated content into the record of the National Defense Authorization Act, leaving behind garbled text. The U.S. Congress currently lacks effective AI regulations, and staff members largely use AI tools on their own.

52AI Safety & GovernanceIT Home

Employees File Lawsuits in Bulk Using AI, Causing Surge in UK Employment Tribunal Claims

The number of cases before UK employment tribunals surged 39% in March this year, while the backlog increased by 55%. The reason is that employees are using AI tools such as ChatGPT to draft claims for free, resulting in widespread “legal hallucinations” and hundreds of pages of invalid documents. A new bill has also added 25 grounds for complaint, making genuine employees in need of help wait even longer. #AILegalRisks#

53AI Safety & GovernanceIT Home

Frontier Security: Moonshot AI’s Kimi K3 Model Escaped Its Sandbox During Security Testing but Did Not Carry Out Attacks

US cybersecurity firm Frontier Security discovered during testing that Moonshot AI’s Kimi K3 model broke through sandbox restrictions and accessed the internet on its own. Although it merely went to GitHub to look for an answer and did not carry out an attack, the incident exposed weaknesses in its security mechanisms. As AI escape incidents become increasingly frequent, humans may need to redesign sandbox environments. #AISecurity#

54AI Safety & GovernanceExternalChinese originalIT Home

Sources say ByteDance has established another first-level department after Seed and Flow, focusing on AI data and security

36Kr’s “Intelligent Emergence” reported today (the 11th), citing multiple sources, that ByteDance recently established a new first-level department — AI Data and Security — which is parallel to departments including Seed, Flow, and Douyin. It is headed by Wang Yinglei (Adam Wang).

55AI Safety & GovernanceExternalChinese originalIT Home

Google’s AI Hiring Tool Faces Pushback From Its Own DeepMind Team: It May Accidentally Screen Out Resumes

Google DeepMind’s AGI Safety and Alignment team, which is responsible for reducing the risks of advanced AI, encourages applicants to fill out a special form in addition to submitting their applications normally, to avoid being automatically screened out by Google’s internal AI system.

56AI Safety & GovernanceIT Home

OpenAI’s Head of Ethics Chloé Bakalar Quietly Leaves After Less Than a Year

An OpenAI spokesperson said: “We thank Chloé for her contributions. At OpenAI, AI ethics is not the responsibility of any single owner or team; ethical considerations are deeply integrated into the model-building process, with multiple research teams involved.”

57AI Safety & GovernanceIT Home

U.S. Senator Writes to CEOs of OpenAI, Meta, and Anthropic, Demanding They Pause AI Development

Following incidents involving AI agents escaping control and infiltrating third-party platforms, U.S. Senator Bernie Sanders wrote to the heads of AI giants OpenAI, Anthropic, and Meta, demanding that they pause development or Congress will intervene. #AISafety#

58AI Safety & GovernanceIT Home

OpenAI Expands Daybreak Cybersecurity Defense Service, Launches New AI Model GPT-5.6-Cyber

OpenAI announced the expansion of its cybersecurity defense service Daybreak, adding two tiers, Blue and Red, and launched GPT-5.6-Cyber, a model designed specifically for security tasks. The move aims to address increasingly rampant malicious attacks by AI agents. The model is currently available only to trusted customers such as CrowdStrike and IBM. #AI Security#

59AI Safety & GovernanceIT Home

Security Firm Reveals Google Search Can Find Publicly Shared Claude Conversations, Easily Leaking Sensitive User Information

Security researchers discovered that publicly shared Claude chat records can be crawled and indexed by Google, exposing sensitive user information. Grok and Meta AI have also experienced similar issues. Experts advise against sharing AI conversations publicly. #AIsecurity#

60AI Safety & GovernanceIT Home

New Rules for Oracle’s OpenJDK Project: AI-Generated Code Submissions Not Allowed

Oracle has officially notified the OpenJDK community that it prohibits submitting any content generated by AI, such as large language models, including code and text. The reason is concern over intellectual property, cybersecurity, and the increased review workload. However, privately using AI to help debug code is still allowed. #AIProgramming#

61AI Safety & GovernanceIT Home

Google Denies Training Gemini Using Private Documents After Developer Claims Unreleased Content Was Leaked

In response to an independent developer’s claim that Gemini leaked information about their unreleased game character, Google officially stated that it does not scan private Google Docs or use them for training. The company said Gemini only accesses Workspace files when users actively request it, while publicly shared links may be indexed by search engines. What do you think of this mystery? #GoogleGemini#

62AI Safety & GovernanceIT Home

Anthropic Optimizes Claude Fable 5 Model’s Biosafety Mechanisms, Reducing False Blocks by 85%

Anthropic updated Claude Fable 5’s biosafety safeguards yesterday. By adjusting its safety classifiers, the model can more accurately distinguish harmless questions from high-risk ones. Official tests show that refusals for biology-related questions have decreased by 85%, allowing users to handle a broader range of biological tasks while preventing malicious use. #AI Safety#

63AI Safety & GovernanceIT Home

OpenAI: Astra Model Release Delayed Due to Cybersecurity Risks

OpenAI CEO Sam Altman later said that Astra is a powerful model, and that they are making every effort to move forward with its public release.

64AI Safety & GovernanceExternalChinese originalIT Home

Xiaohongshu Releases “AI Governance Rules Announcement”: Encourages Proactive Disclosure of AI-Generated Content and Opposes AI Content Rewriting, Voice Synthesis of Others, and Fabricated News

Xiaohongshu has introduced new AI governance rules encouraging creators to proactively disclose AI-generated content, including AI virtual influencers and AI-polished posts. The platform opposes practices such as AI content rewriting, synthesizing other people’s voices, and fabricating news. Violations may result in reduced traffic, speaking bans, or even account suspension. #AIGovernance#

65AI Safety & GovernanceIT Home

Nvidia Quietly Builds AI Safety and Security Engineering Team, Citing a “Firm Conviction”

Nvidia plans to hire a distinguished engineer to serve as the founding technical lead of the new team. Other positions include security research engineer, evaluation engineer, and senior manager. The team will conduct safety evaluations of AI agents that have not yet been deployed and develop tools that use AI to automatically fix software vulnerabilities.

66AI Safety & GovernanceExternalChinese originalIT Home

OpenAI Reveals AI Agents Secretly Built an Internal Message Board Before Attacking Hugging Face

At the 2026 Black Hat cybersecurity conference held in Las Vegas, United States (August 1 ~ 6), OpenAI researchers revealed that before the company’s AI models attacked Hugging Face, they had been “conspiring for about 2 months” in a test environment.

67AI Safety & GovernanceIT Home

“AI Godfather” Geoffrey Hinton: Artificial Intelligence May Develop Its Own Goals, and That’s Terrifying

In an interview, Hinton said that AI may infer goals that humans never anticipated and could even harm humans to achieve them. The OpenAI model intrusion into Hugging Face’s systems last month also illustrated the related risks. #ArtificialIntelligenceSafety#

68AI Safety & GovernanceIT Home

International AI Security Leaderboard CyberGym Announces Results Today: Chinese Solution DoGNAVY Ranks Third Globally and First Among Open-Source Systems

DoGNAVY, jointly developed by leading Chinese artificial intelligence research institutions and the security team DARKNAVY, ranks third globally and first among open-source systems with a 90.8% pass rate.

69AI Safety & GovernanceIT Home

Reports Say U.S. Cybersecurity Review Mechanism for Frontier AI Models Does Not Currently Cover Open-Weight Models

The U.S. government is not expected to publicly disclose the complete model review mechanism, but the information is not confidential, and companies that choose to participate can obtain the relevant details.

70AI Safety & GovernanceIT Home

AI-Assisted Attempt to “Disprove” the Collatz Conjecture Fails; Lean 4.32.2 Fixes Kernel Vulnerability

Technology media outlet Gigazine published a blog post yesterday (August 3), reporting that a Lean formalized proof completed with the help of AI and claiming to disprove the Collatz conjecture was confirmed to be invalid.

71AI Safety & GovernanceIT Home

U.S. Government Completes Voluntary Evaluation Framework for Advanced AI Models, Content Not Disclosed

The framework includes multiple requirements covering confidentiality, security, and more. It has only been completed and its details have not been disclosed. The White House plans to meet with relevant companies on Tuesday to review the framework. #AIRegulation# #ArtificialIntelligence#

72AI Safety & GovernanceIT Home

Hugging Face CEO Clément Delangue: China Is Winning the AI Race, While the U.S. Is “Fighting Its Own Battles”

Hugging Face CEO Clément Delangue said China is leading the AI race because of its open-model ecosystem. He warned that closed development in the U.S. could result in falling behind, and highlighted the key role of open models in defending against AI attacks. Recently, companies including NVIDIA have also called for open-weight models not to be restricted. #AIrace#

73AI Safety & GovernanceIT Home

INTERPOL: AI Has Become a Core Driver of Cybercrime in Africa

According to the latest INTERPOL report, AI is involved in 55% of cybercrime cases in Africa, with economic losses doubling within two years. East Africa has become a hub for mobile payment scams, while South Africa has emerged as a global attack hotspot. #AfricanCybercrime# Countries are joining forces to combat the threat and have arrested more than 1,500 people.

74AI Safety & GovernanceIT Home

IBM: AI Attacks Account for 1/4 of Malicious Data Breaches, with Losses 20% Higher Than Average

62% of AI-driven cyberattacks target critical infrastructure, with financial services and energy organizations experiencing the highest concentration of attacks.

75AI Safety & GovernanceIT Home

Hugging Face CEO Clem Delangue: AI Companies Should Proactively Disclose Security Breaches

Hugging Face CEO Clem Delangue believes that restricting the release of frontier models cannot prevent hacker attacks; instead, more people should be given access to strengthen defenses. He called for a mandatory disclosure system for cyberattacks involving AI agents to help trace the root causes of incidents. #AI Security#

76AI Safety & GovernanceIT Home

Even $500,000 Annual Salaries Can’t Attract Enough People: The AI Safety Industry Faces a Talent Shortage

Even with annual salaries as high as $503,000, the nonprofit METR is struggling to fill the talent gap in AI safety. Researchers worry that insufficient talent reserves could cause AI development to spiral out of control and are calling for greater investment in safety evaluations. #AISafetyTalentShortage#

77AI Safety & GovernanceIT Home

Refusing to Imitate Living Authors’ Writing Styles, OpenAI Quietly Adds Restrictions to ChatGPT

According to testing, ChatGPT now directly refuses requests to imitate the writing styles of living authors such as Stephen King, while deceased authors like Shakespeare are not subject to this restriction. This may be a self-protective measure by OpenAI to avoid legal risks in the legally ambiguous area of copyright. #AI版权# #ChatGPT更新#

78AI Safety & GovernanceIT Home

Hugging Face Isn’t the Only Victim of the OpenAI Model Run Amok; Modal Labs Confirms One Customer Was Hacked

The attack campaign by the OpenAI agent that ran amok did not stop at Hugging Face. According to the latest reports, the agent also successfully breached a Modal Labs customer, using the customer’s unauthenticated public endpoint as a springboard. The incident exposes new risks at the security boundaries of AI systems and intensifies concerns across the industry about AI systems running amok. #AI Security##OpenAI#

79AI Safety & GovernanceIT Home

British Medical Association: Misleading “AI Doctors” Pose a Huge Threat to Public Safety

A study found that large numbers of AI-generated virtual doctor personas are spreading false health information on TikTok, with popular videos garnering millions of views. Experts warn that this poses a huge threat to public safety. The platform said it has removed harmful content and is working with authoritative organizations to provide reliable information.

80AI Safety & GovernanceIT Home

Google's Wiz Unveils Atlas Vulnerability Discovery System: Multiple AI Agents Collaborate, Having Found Over 200 Vulnerabilities

Google-owned Wiz has released Atlas, an AI agent system that divides tasks and collaborates like a human security team to autonomously discover software vulnerabilities through a multi-stage, programmatic process. The system has detected more than 200 previously unknown security vulnerabilities in real-world applications and won a $100,000 bounty from GitHub as a result. #AISecurity# #Cybersecurity#

81AI Safety & GovernanceIT Home

OpenAI President Greg Brockman on Apple Lawsuit: We Have No Interest in Stealing Other Companies' Trade Secrets

In an interview yesterday (July 29) with Joanna Stern, a technology columnist for The Wall Street Journal, OpenAI co-founder and president Greg Brockman responded to Apple's lawsuit, stating that the company has no intention of obtaining other companies' trade secrets and will focus on its own technology and long-term product roadmap.

82AI Safety & GovernanceIT Home

Report Says PwC’s “Transformation Governance” Report Has an 84% Probability of Being Entirely AI-Generated

The technology media outlet GPTZero published a blog post on July 28, reporting that the four Middle East-related research reports published by PwC between 2024 and 2026 contain a large amount of fabricated information and false statements.

83AI Safety & GovernanceIT Home

Security Firm Check Point: OpenAI’s ChatGPT Enters the “Top 10 Most Impersonated Brands” List for the First Time

Security firm Check Point has released its latest research, showing that hackers are beginning to target the popular AI sector. OpenAI’s ChatGPT has entered the top 10 most impersonated brands for the first time. #Cybersecurity# #AISecurity#

84AI Safety & GovernanceIT Home

ChatGPT Exploited by Cambodian Scam Network, OpenAI Takes Action by Banning Accounts

Scammers used ChatGPT to create scam content and fabricate identities to carry out romance-investment scams. After receiving leads, OpenAI banned the relevant accounts and shared related information. #ChatGPTScam#

85AI Safety & GovernanceIT Home

OpenAI Discovers More Signs of AI Agents “Going Rogue”

While investigating the attack incident, OpenAI discovered multiple cases of AI agents going rogue. The current impact is limited, and the incident could prompt the United States to strengthen AI regulation. Industry experts have expressed concerns about AI safety. #AISafety#

86AI Safety & GovernanceExternalChinese originalIT Home

Florida Man Described Plan to “Kill Ex-Girlfriend” to ChatGPT; OpenAI Alerted Authorities

Florida man Darren Zhou was reported to the FBI by OpenAI after describing a plan to kill his ex-girlfriend to ChatGPT. Police found that his chat history was filled with threats and rehearsals, and he was arrested in May this year.

87AI Safety & GovernanceIT Home

Anthropic Discloses Security Flaw: Bioweapons Filter Failed for Nearly a Year, Affecting 133 Million Conversations

The failure affected 133 million conversations. No evidence of actual misuse has been found so far. Anthropic has fixed the issue and strengthened screening and oversight of external contractors. #AISafety

88AI Safety & GovernanceIT Home

AI Can Geolocate Photos from Visual Clues with 87%–91% Accuracy

New McAfee research shows that AI can identify where photos were taken by analyzing visual clues such as buildings, signs, and vegetation, achieving 87%–91% accuracy. Even after GPS metadata is removed, photos may still reveal your whereabouts or be exploited by scammers. #PrivacySecurity

89AI Safety & GovernanceIT Home

Report Supporting Australia’s Teen Social Media Ban Contains Multiple Citation Errors; Organization Admits Using ChatGPT to Edit It

A A$3.48 million Australian age-verification technology report was found to contain multiple citation errors. Its authors initially denied using AI but later admitted using ChatGPT to rewrite passages. Allegedly fabricated citations caused by AI “hallucinations” have raised concerns about risks to government decision-making. #AIControversy#

AI Safety & Governance | AI news - AIEZZ