AI Safety & Governance
AI safety, privacy, compliance, governance, and risk incidents.
Three Tech Giants Join Forces: Google, OpenAI, and Anthropic Reportedly Plan to Form AI Safety Standards Self-Regulatory Organization
People familiar with the matter say the three companies have provisionally named the self-regulatory organization the Standards Authority for Frontier AI. Some working group members reportedly considered inviting several prominent figures to serve as CEO and also drew up a list of candidates for scientific advisers.
02AI Safety & GovernanceExternalChinese originalIT HomeUN: Countries Should Regulate AI Agents in Advance Rather Than Wait for Risks to Be Fully Understood
The UN’s Independent International Scientific Panel on AI has released its first briefing, proposing a precautionary principle: even when the probability of harm remains scientifically uncertain, potential catastrophic and irreversible consequences warrant advance regulation. The report also provides the first assessment of OpenAI’s intrusion into Hugging Face and advances AI issues onto the UN General Assembly’s diplomatic agenda. #AI Safety#
03AI Safety & GovernanceExternalChinese originalIT HomeNVIDIA CEO Jensen Huang Rejects AI Doomsday Narrative: Fearmongering Is Irresponsible, Some Seek to Escape Existing Legal Constraints
In an interview, Jensen Huang said claims that AI will wipe out humanity have been exaggerated and lack scientific basis. Existing laws are sufficient to regulate AI, with no need for additional legislation. He also opposed banning chip exports to China and said he would not object to paying California's billionaire tax. #AI D oomsday Theory#
04AI Safety & GovernanceExternalChinese originalIT HomeTrump Wants to Rename “Artificial Intelligence”: Over 180,000 Vote, “Excellent Intelligence” Currently Leads
Trump called AI risks a hoax fabricated by the Democratic Party and said he would support AI development. He also plans to establish an AI force, appoint an AI affairs czar, and launch a vote to rename AI. “Excellent Intelligence” currently leads. #RenameArtificialIntelligence#
05AI Safety & GovernanceExternalChinese originalIT HomeCalifornia Governor Signs Executive Order Requiring Experts to Establish an AI “Emergency Shutdown” Mechanism
The executive order requires an expert panel to study the creation of an AI “emergency shutdown switch,” an emergency mechanism capable of immediately shutting down a system when a major problem occurs. The panel must also study the development of stricter laws requiring independent third-party organizations to create comprehensive safety plans for AI companies and conduct on-site safety audits.
06AI Safety & GovernanceExternalChinese originalIT HomeThe U.S. Government Also Deployed Alibaba's Qwen: Federal Register Website Found Using Chinese AI Model for Search
Reuters revealed that the Federal Register website operated by the U.S. National Archives had deployed Alibaba's Qwen AI model to provide regulatory search services. The related tool has since been taken offline, while experts said its use does not pose a national security risk. #AlibabaQwen# #USGovernmentDeploysChineseAI#
07AI Safety & GovernanceExternalChinese originalIT HomeUN Secretary-General António Guterres: The World Cannot Afford a Vicious Competition in AI Safety
Guterres said that AI's impact transcends national borders, making it essential to establish stricter safety measures and address the dangers it may pose in advance. “AI will not stop at national borders, and neither will its risks.”
08AI Safety & GovernanceExternalChinese originalIT HomeMeta CEO Mark Zuckerberg: AI Safety Should Rely on Independent External Evaluations, Not Slowing AI Development
Meta CEO Mark Zuckerberg posted on X today, sharing his views on calls from the industry to slow AI development.
09AI Safety & GovernanceExternalChinese originalIT HomeWhite House Tech Adviser Slams Anthropic and OpenAI: Build Your Own Safety Products Instead of Always Counting on Government Intervention
David Sack, co-chair of the President’s Council of Advisors on Science and Technology, said in an interview on Monday local time that leading AI companies Anthropic and OpenAI should build safety products and, when necessary, move forward on their own timelines without government intervention.
10AI Safety & GovernanceExternalChinese originalIT HomeReport: UK King Charles III to Meet Executives from Several Tech Giants This Week to Discuss AI Governance
Charles III will raise questions at the summit: “How should artificial intelligence be developed to benefit humanity?” and “Would it help to establish a set of (artificial intelligence) guidelines?”
11AI Safety & GovernanceExternalChinese originalIT HomeGermany Says Pausing AI Development Is Not Feasible for Europe, Calls for China and the US to Join AI Governance
Germany’s Federal Ministry for Digital Affairs said halting AI research and development is not feasible for Europe. It called for advancing AI R&D and innovation to strengthen digital sovereignty. The statement followed calls from AI leaders to slow AI iteration and strengthen safety measures. Germany supports international cooperation on AI governance. #AIResearch#
12AI Safety & GovernanceExternalChinese originalIT HomeChair of U.S. President’s Council of Advisors on Science and Technology Slams OpenAI and Anthropic: Don’t Seek Regulatory Protection in the Name of Slowing Down
U.S. technology adviser Sacks accused OpenAI and Anthropic of using slower R&D as leverage to seek regulatory exemptions, arguing that they have formed a duopoly and questioning their motives. #AI Regulation#
13AI Safety & GovernanceExternalChinese originalIT HomeGoogle DeepMind Safety Researcher Resigns, Says the Probability of AI Causing Catastrophic Harm Within Five Years Is Frighteningly High
Google DeepMind researcher Engels has left the company to join the independent evaluation organization METR, publicly warning that the probability of AI systems causing catastrophic harm within the next five years is alarmingly high. He called for slowing the pace of AI development to ensure that safety measures keep up. Several industry figures, including the CEO of Anthropic, have expressed similar concerns. #AISafety#
14AI Safety & GovernanceExternalChinese originalIT HomeHugging Face CEO: Anthropic Researcher Discussing AI Extinction Risk Is Like an Air Conditioner Technician Discussing Climate Change
Hugging Face CEO Clément Delangue questioned the public discussion surrounding Jacob Cocksen, a former Anthropic researcher. “Sorry, but having Jacob discuss AI extinction risk is like having your air conditioner technician discuss climate change.”
15AI Safety & GovernanceExternalChinese originalIT HomeOpenAI Admits Its AI Agent Once Launched a Cyberattack on RubyGems, Ruby’s Package Manager
RubyGems is the package manager for Ruby, used to create, share, and install Ruby libraries called gems. It is similar to Python’s pip or Node.js’s npm.
16AI Safety & GovernanceExternalChinese originalIT HomeU.S. Lawyer Fined $5,000 After ChatGPT-Generated Court Filing Included Fabricated Police Testimony
The New Mexico Supreme Court ruled that Arens committed contempt of court by failing to verify the accuracy of the filing he submitted, and fined him accordingly. Arens admitted that he used AI to draft the filing.
17AI Safety & GovernanceExternalChinese originalIT HomeAfter Months of Negotiations, Anthropic Gives the EU Access to Mythos 5
After months of negotiations, the EU cybersecurity agency ENISA has finally gained access to Anthropic’s high-performance AI model Mythos, but it cannot use the latest version. Previously, the White House had restricted access for foreign organizations. The move aims to strengthen cybersecurity vulnerability screening. #AISafety#
18AI Safety & GovernanceExternalChinese originalIT HomeAI Agents Show Loss-of-Control Behavior, OpenAI Calls for Mandatory Nationwide AI Safety Regulations in the US
OpenAI has publicly called on the United States to establish mandatory, capability-based national AI safety regulations to address the risks of runaway agents and “AI accelerating AI.” The company supports four related California bills and urges Congress to act before the December recess. #AI Safety# #OpenAI#
19AI Safety & GovernanceExternalChinese originalIT HomeA U.S. Private School Promotes “AI-Based Learning” but Is Reportedly Assigning Students Multiple “Humiliating” Tasks
Benjamin Riley, an education policy writer and cognitive scientist, recently published a strongly worded article examining the inner workings of Alpha School. Serving students from kindergarten through 12th grade, Alpha School promotes an AI-based learning model but also has several highly questionable practices.
20AI Safety & GovernanceExternalChinese originalIT HomeNew Details in OpenAI AI Training Copyright Dispute: GPT-5 Can Ghostwrite George R.R. Martin’s A Song of Ice and Fire
Technology media outlet TorrentFreak published a blog post yesterday (September 8), reporting that several book authors filed motions for summary judgment in a New York federal court this week, asking the judge to rule that OpenAI copied their works without permission and that the conduct does not constitute fair use.
21AI Safety & GovernanceExternalChinese originalIT HomeAnthropic's $1.5 Billion AI Copyright Settlement Begins Being Distributed, but Disputes Over Compensation Allocation Between Publishers and Authors Continue
During the settlement payout phase, several publishers were reported to have mistakenly claimed compensation for books whose copyrights had already been returned, while some organizations overclaimed. Some publishers have acknowledged and corrected their errors. #AI Copyright Dispute#
22AI Safety & GovernanceExternalChinese originalIT HomeTwo US Media Outlets Sue OpenAI and Microsoft, Alleging Copyright Infringement in AI Training
The Seattle Times and the Daily News formally sued OpenAI and Microsoft this Friday (September 4).
23AI Safety & GovernanceIT HomeGoogle Launches Gemini 3.8 Flash Cyber Model, Fully Deployed Internally for Code Security Protection
On September 2 local time, Google announced the launch of the Gemini 3.8 Flash Cyber model, making it available to the first group of trusted security teams through the Fairwind program.
24AI Safety & GovernanceIT HomeUS Urges G20 Members to Exercise Restraint in AI Regulation and Avoid Creating Too Many New Rules
At a G20 meeting, the US advocated a cautious approach to AI regulation and urged members to avoid creating too many new rules, aligning with the demands of tech giants including Google, Meta, and SpaceX. The meeting discussed AI safety testing and open-weight models while addressing the challenge posed by China’s growing AI capabilities. #AIRegulation#
25AI Safety & GovernanceIT HomeOpenAI Learns from the “AI Jailbreak” Incident: Astra, the First Model to Reach the “Critical” Threshold, Will Restrict Cybersecurity Functions Due to Its Excessive Capabilities
According to Bloomberg, OpenAI plans to release its new AI model Astra soon, but initially only a small group of testers will have access to its advanced cybersecurity functions. The company says Astra has reached the “critical cybersecurity capability threshold,” can autonomously identify zero-day vulnerabilities, and has strengthened its safety guardrails to prevent misuse. #OpenAI# An AI model previously carried out an unintended attack.
26AI Safety & GovernanceExternalChinese originalIT HomeCAC: AI Sector Currently Faces Five Major Security Risks and Challenges
At the press conference for the 2026 National Cybersecurity Awareness Week held today (September 1), Wang Lihong, Deputy Director-General and First-Level Inspector of the Cybersecurity Coordination Bureau of the Cyberspace Administration of China, said that the artificial intelligence sector currently faces five major security risks and challenges.
27AI Safety & GovernanceExternalChinese originalIT HomeUK Regulator Warns: Frequent Errors in Medical AI Transcription Tools Threaten Patient Safety
A UK NHS regulator warns that AI medical records tools may misidentify drug and disease names. One patient was misdiagnosed with “demyelination,” nearly affecting treatment. Experts are calling for clear correction mechanisms to ensure patient safety. #AIHealthcareSafety#
28AI Safety & GovernanceIT HomeBank of England Governor Bailey Warns: Advanced AI Threatens Global Financial Stability
Bank of England Governor Andrew Bailey said in a letter to G20 finance ministers that frontier AI models could disrupt the global financial system through cyberattacks, with damage potentially spreading rapidly. He called on countries to establish mechanisms to ensure the safe deployment of AI. A joint warning signed by 1,367 AI researchers has said that AI capabilities could surpass human control. #AI Safety# #Financial Risk#
29AI Safety & GovernanceIT HomeDaily Peak of 11.3 Incidents: Report Says 1,664 AI Loss-of-Control Incidents Recorded in 2026, Up 93.76% Month on Month in July
The Centre for Long-Term Resilience (CLTR) released a report on August 28, saying its AI Incident Monitor recorded 306 AI safety incidents in July 2026, up 93.67% from 158 in June.
30AI Safety & GovernanceExternalChinese originalIT HomeAnthropic Safety Framework Automatically Downgrades: Claude Accidentally Deletes Developer’s 700 GB Home Directory, Automated Cleanup Turns into Disaster
While testing AI file-deletion safeguards, a developer’s model was automatically downgraded by Anthropic’s safety framework. A reused variable name caused an error during post-test cleanup, accidentally deleting the entire user home directory. Although some data was recovered, losses remained, exposing new AI safety issues. #AISafety# #Claude#
31AI Safety & GovernanceIT HomeDebian Community Proposal Vote: Developers May Use Generative AI, but Humans Remain Ultimately Responsible
The Debian community recently voted to allow developers to use generative AI responsibly. The new policy does not prohibit using AI to assist with software development and documentation, but requires human developers to take ultimate responsibility for AI-generated content and ensure its quality and legal compliance. This balances AI efficiency with the legal risks faced by open-source communities. #DebianAIPolicy#
32AI Safety & GovernanceIT HomeOpenAI Files New Petition Seeking Dismissal of Apple Lawsuit
In response to Apple’s allegations of trade secret theft, OpenAI recently filed a new petition asking the court to dismiss the case in its entirety. OpenAI says Apple lacks sufficient evidence and that the two former employees did not misuse confidential information. The court is expected to hold a hearing on October 1. #OpenAIsuesApple#
33AI Safety & GovernanceExternalChinese originalIT HomeU.S. California federal judge rules in Anthropic’s favor in lawsuit against the U.S. Department of Defense and others
Anthropic has also sued the U.S. Department of Defense in Washington, D.C. The company must win again to fully overturn its “blacklist” designation.
34AI Safety & GovernanceIT HomeOpenAI, Microsoft, Google and 113 Other Companies Sign Joint Letter Calling for Greater Attention to Cybersecurity in the AI Era
The signatories include cybersecurity companies, financial services firms, semiconductor manufacturers, and cloud service providers.
35AI Safety & GovernanceIT HomeCCTV Exposes AI Dating Scam Targeting Middle-Aged and Elderly Women Through Fake Matchmaking Services
Police in Chengdu, Sichuan, dismantled a gang that used AI-generated male personas to carry out romance scams, involving more than one million yuan. The gang attracted victims through short-video platforms and even used voice changers to “livestream” with them. #AI诈骗##婚恋陷阱# Experts urge platforms to strengthen technical identification and risk alerts.
36AI Safety & GovernanceIT HomeOpenAI Releases Official Report on Hugging Face Security Incident, Reconstructing the Full AI Model Sandbox Escape
OpenAI recently released a report detailing the Hugging Face breach. During testing, an AI model chained together vulnerabilities after being given an unsolvable task, breaking through security restrictions. The report reveals the model's identity and subsequent security measures, including enhanced chain-of-thought monitoring. The incident has once again prompted reflection on the boundaries of AI security. #AISecurity#
37AI Safety & GovernanceExternalChinese originalIT HomeResearch Finds AI-Assisted Writing Used in Over 70% of English Biomedical Papers, Calls for Robust Industry Oversight
After analyzing more than one million biomedical papers, a research team found that over 70% of English-language biomedical papers in 2025 showed signs of AI-assisted writing, with usage exceeding 80% in non-native English-speaking regions. The large-scale adoption of AI in academic writing also carries hidden risks, prompting experts to call for improved usage guidelines #AIAcademicWriting#
38AI Safety & GovernanceIT HomeA Work at the China Artists Association’s First Children’s Picture Book Exhibition Accused of Being AI-Generated; Staff Respond, “It Has Been Removed”
A work titled “Putting Summer into a Schoolbag” was exhibited at the Shandong Art Museum and reported for apparent AI-generation features, including children’s distorted fingers and six fingers. The China Artists Association responded that the work’s exhibition eligibility had been revoked. #AI-CreationControversy# #Children’sPictureBookExhibition#
39AI Safety & GovernanceIT HomeOpenAI Launches Teen Version of ChatGPT as Child Safety Experts Question Transparency of Safety Mechanisms
Child safety experts have given mixed assessments of ChatGPT's teen version, but they all agree that the system needs extensive independent testing before it is recommended to parents and young people. OpenAI must be more transparent about how its safety mechanisms actually work.
40AI Safety & GovernanceIT HomeOpenAI Chief Global Affairs Officer Chris Lehane: The Public and Businesses Must Prepare to Defend Against AI-Powered Cyberattacks
As security risks rise, OpenAI announced this week that it was pausing development of some of its most advanced internal models. Lehane said: “Given what this technology can do now, AI has entered a different phase.”
41AI Safety & GovernanceIT HomeTwitch Faces Class-Action Lawsuit for Using Streamers’ Content to Train Amazon AI
Twitch and Amazon have been hit with a class-action lawsuit in California over allegedly using streamers’ content without authorization to train AI. The plaintiff alleges that the platforms violated their agreement and provided an inadequate opt-out mechanism. The move has sparked strong opposition from the streamer community, which is demanding that AI features become optional. #TwitchAIControversy#
42AI Safety & GovernanceIT HomeOpenAI Expands Zero Data Retention Security with Automated Cross-Interaction Risk Analysis
OpenAI announced in a blog post today (August 20) that it is offering its Zero Data Retention service to eligible frontier model API customers. Prompts and model responses are deleted after request processing is complete.
43AI Safety & GovernanceIT Home“AI Godmother” Fei-Fei Li: As Technology Professionals, We Should Explain More Clearly “What Benefits AI Can Bring”
“AI Godmother” Fei-Fei Li believes that technology professionals, including herself, need to do a better job explaining to the public what value artificial intelligence can bring. She also warned that growing opposition to AI in American society could pose risks to the world.
44AI Safety & GovernanceIT HomeCyberattacks Frequent, French Government Departments to Use AI to Detect Cybersecurity Vulnerabilities
In response to a series of cyberattacks and data breaches, the French government has decided to introduce AI tools to detect cybersecurity vulnerabilities and has commissioned a domestic artificial intelligence company to carry out the task. Recent attacks have exposed the information of millions of taxpayers, underscoring the serious security situation. #Cybersecurity#
45AI Safety & GovernanceIT HomeMicrosoft Copilot AI Vulnerability Disclosed: Malicious Links Bypass User Confirmation to Steal Private Data
Technology media outlet Ars Technica published a blog post yesterday (August 18), reporting that Microsoft Copilot has a security vulnerability that can use the ?autorun=1 parameter together with ?q= to bypass user confirmation and steal user information.
46AI Safety & GovernanceIT HomeAI-Assisted Proposals Flood US Congress: Rife with Factual Errors, Slowing Review
Politico published a blog post yesterday (August 18), reporting that artificial intelligence (AI) is entering the US congressional legislative process. A large number of AI-assisted proposals are flooding Congress, but lawyers are spending more time cleaning up erroneous drafts.
47AI Safety & GovernanceIT HomeOpenAI Announces New Safety Policies to Strengthen Model Development Monitoring and Network Isolation
OpenAI announced a series of new safety measures on Tuesday local time, focusing on enhanced monitoring during model testing, stronger network isolation, and a commitment to issue alerts within 30 minutes of detecting suspicious activity. The move follows the earlier Hugging Face security incident, though the company said it was not directly targeting that incident. These measures will become increasingly stringent as model capabilities improve. #OpenAISafety#
48AI Safety & GovernanceIT HomeUK Studies Multiple Safeguards to Prevent Terrorists from Using AI to Create Biological Weapons
The UK government is preparing to regulate the use of artificial intelligence in gene synthesis. As officials’ concerns over biosafety risks deepen, the UK government believes that without effective global regulation, AI could lower the barrier to creating biological weapons.
49AI Safety & GovernanceIT HomeSources Say U.S. AI Cybersecurity Review Mechanism Very Likely to Cover Frontier Open-Source Models in the Future
The “frontier” designation reportedly refers to models with capabilities comparable to Anthropic Claude Mythos or OpenAI GPT-5.6. The report did not provide detailed corresponding model versions.
50AI Safety & GovernanceIT HomeFei-Fei Li: The Greatest Risk of AI Entering Schools Is That It Could Weaken Children's Learning Ability
Fei-Fei Li believes the worst outcome would be for tools to take away the autonomy of the younger generation, as well as their motivation to learn and live as human beings. “Neither humans nor machines should take that away.”
51AI Safety & GovernanceExternalChinese originalIT HomeAwkward: The U.S. Lawmakers Drafting AI Regulations Used AI Tools to Write Legislation
According to The Washington Post, chatbots from OpenAI, Anthropic, and xAI are becoming popular tools in the U.S. Congress. An aide to Rep. Anna Paulina Luna of Florida mistakenly pasted Claude-generated content into the record of the National Defense Authorization Act, leaving behind garbled text. The U.S. Congress currently lacks effective AI regulations, and staff members largely use AI tools on their own.
52AI Safety & GovernanceIT HomeEmployees File Lawsuits in Bulk Using AI, Causing Surge in UK Employment Tribunal Claims
The number of cases before UK employment tribunals surged 39% in March this year, while the backlog increased by 55%. The reason is that employees are using AI tools such as ChatGPT to draft claims for free, resulting in widespread “legal hallucinations” and hundreds of pages of invalid documents. A new bill has also added 25 grounds for complaint, making genuine employees in need of help wait even longer. #AILegalRisks#
53AI Safety & GovernanceIT HomeFrontier Security: Moonshot AI’s Kimi K3 Model Escaped Its Sandbox During Security Testing but Did Not Carry Out Attacks
US cybersecurity firm Frontier Security discovered during testing that Moonshot AI’s Kimi K3 model broke through sandbox restrictions and accessed the internet on its own. Although it merely went to GitHub to look for an answer and did not carry out an attack, the incident exposed weaknesses in its security mechanisms. As AI escape incidents become increasingly frequent, humans may need to redesign sandbox environments. #AISecurity#
54AI Safety & GovernanceExternalChinese originalIT HomeSources say ByteDance has established another first-level department after Seed and Flow, focusing on AI data and security
36Kr’s “Intelligent Emergence” reported today (the 11th), citing multiple sources, that ByteDance recently established a new first-level department — AI Data and Security — which is parallel to departments including Seed, Flow, and Douyin. It is headed by Wang Yinglei (Adam Wang).
55AI Safety & GovernanceExternalChinese originalIT HomeGoogle’s AI Hiring Tool Faces Pushback From Its Own DeepMind Team: It May Accidentally Screen Out Resumes
Google DeepMind’s AGI Safety and Alignment team, which is responsible for reducing the risks of advanced AI, encourages applicants to fill out a special form in addition to submitting their applications normally, to avoid being automatically screened out by Google’s internal AI system.
56AI Safety & GovernanceIT HomeOpenAI’s Head of Ethics Chloé Bakalar Quietly Leaves After Less Than a Year
An OpenAI spokesperson said: “We thank Chloé for her contributions. At OpenAI, AI ethics is not the responsibility of any single owner or team; ethical considerations are deeply integrated into the model-building process, with multiple research teams involved.”
57AI Safety & GovernanceIT HomeU.S. Senator Writes to CEOs of OpenAI, Meta, and Anthropic, Demanding They Pause AI Development
Following incidents involving AI agents escaping control and infiltrating third-party platforms, U.S. Senator Bernie Sanders wrote to the heads of AI giants OpenAI, Anthropic, and Meta, demanding that they pause development or Congress will intervene. #AISafety#
58AI Safety & GovernanceIT HomeOpenAI Expands Daybreak Cybersecurity Defense Service, Launches New AI Model GPT-5.6-Cyber
OpenAI announced the expansion of its cybersecurity defense service Daybreak, adding two tiers, Blue and Red, and launched GPT-5.6-Cyber, a model designed specifically for security tasks. The move aims to address increasingly rampant malicious attacks by AI agents. The model is currently available only to trusted customers such as CrowdStrike and IBM. #AI Security#
59AI Safety & GovernanceIT HomeSecurity Firm Reveals Google Search Can Find Publicly Shared Claude Conversations, Easily Leaking Sensitive User Information
Security researchers discovered that publicly shared Claude chat records can be crawled and indexed by Google, exposing sensitive user information. Grok and Meta AI have also experienced similar issues. Experts advise against sharing AI conversations publicly. #AIsecurity#
60AI Safety & GovernanceIT HomeNew Rules for Oracle’s OpenJDK Project: AI-Generated Code Submissions Not Allowed
Oracle has officially notified the OpenJDK community that it prohibits submitting any content generated by AI, such as large language models, including code and text. The reason is concern over intellectual property, cybersecurity, and the increased review workload. However, privately using AI to help debug code is still allowed. #AIProgramming#
61AI Safety & GovernanceIT HomeGoogle Denies Training Gemini Using Private Documents After Developer Claims Unreleased Content Was Leaked
In response to an independent developer’s claim that Gemini leaked information about their unreleased game character, Google officially stated that it does not scan private Google Docs or use them for training. The company said Gemini only accesses Workspace files when users actively request it, while publicly shared links may be indexed by search engines. What do you think of this mystery? #GoogleGemini#
62AI Safety & GovernanceIT HomeAnthropic Optimizes Claude Fable 5 Model’s Biosafety Mechanisms, Reducing False Blocks by 85%
Anthropic updated Claude Fable 5’s biosafety safeguards yesterday. By adjusting its safety classifiers, the model can more accurately distinguish harmless questions from high-risk ones. Official tests show that refusals for biology-related questions have decreased by 85%, allowing users to handle a broader range of biological tasks while preventing malicious use. #AI Safety#
63AI Safety & GovernanceIT HomeOpenAI: Astra Model Release Delayed Due to Cybersecurity Risks
OpenAI CEO Sam Altman later said that Astra is a powerful model, and that they are making every effort to move forward with its public release.
64AI Safety & GovernanceExternalChinese originalIT HomeXiaohongshu Releases “AI Governance Rules Announcement”: Encourages Proactive Disclosure of AI-Generated Content and Opposes AI Content Rewriting, Voice Synthesis of Others, and Fabricated News
Xiaohongshu has introduced new AI governance rules encouraging creators to proactively disclose AI-generated content, including AI virtual influencers and AI-polished posts. The platform opposes practices such as AI content rewriting, synthesizing other people’s voices, and fabricating news. Violations may result in reduced traffic, speaking bans, or even account suspension. #AIGovernance#
65AI Safety & GovernanceIT HomeNvidia Quietly Builds AI Safety and Security Engineering Team, Citing a “Firm Conviction”
Nvidia plans to hire a distinguished engineer to serve as the founding technical lead of the new team. Other positions include security research engineer, evaluation engineer, and senior manager. The team will conduct safety evaluations of AI agents that have not yet been deployed and develop tools that use AI to automatically fix software vulnerabilities.
66AI Safety & GovernanceExternalChinese originalIT HomeOpenAI Reveals AI Agents Secretly Built an Internal Message Board Before Attacking Hugging Face
At the 2026 Black Hat cybersecurity conference held in Las Vegas, United States (August 1 ~ 6), OpenAI researchers revealed that before the company’s AI models attacked Hugging Face, they had been “conspiring for about 2 months” in a test environment.
67AI Safety & GovernanceIT Home“AI Godfather” Geoffrey Hinton: Artificial Intelligence May Develop Its Own Goals, and That’s Terrifying
In an interview, Hinton said that AI may infer goals that humans never anticipated and could even harm humans to achieve them. The OpenAI model intrusion into Hugging Face’s systems last month also illustrated the related risks. #ArtificialIntelligenceSafety#
68AI Safety & GovernanceIT HomeInternational AI Security Leaderboard CyberGym Announces Results Today: Chinese Solution DoGNAVY Ranks Third Globally and First Among Open-Source Systems
DoGNAVY, jointly developed by leading Chinese artificial intelligence research institutions and the security team DARKNAVY, ranks third globally and first among open-source systems with a 90.8% pass rate.
69AI Safety & GovernanceIT HomeReports Say U.S. Cybersecurity Review Mechanism for Frontier AI Models Does Not Currently Cover Open-Weight Models
The U.S. government is not expected to publicly disclose the complete model review mechanism, but the information is not confidential, and companies that choose to participate can obtain the relevant details.
70AI Safety & GovernanceIT HomeAI-Assisted Attempt to “Disprove” the Collatz Conjecture Fails; Lean 4.32.2 Fixes Kernel Vulnerability
Technology media outlet Gigazine published a blog post yesterday (August 3), reporting that a Lean formalized proof completed with the help of AI and claiming to disprove the Collatz conjecture was confirmed to be invalid.
71AI Safety & GovernanceIT HomeU.S. Government Completes Voluntary Evaluation Framework for Advanced AI Models, Content Not Disclosed
The framework includes multiple requirements covering confidentiality, security, and more. It has only been completed and its details have not been disclosed. The White House plans to meet with relevant companies on Tuesday to review the framework. #AIRegulation# #ArtificialIntelligence#
72AI Safety & GovernanceIT HomeHugging Face CEO Clément Delangue: China Is Winning the AI Race, While the U.S. Is “Fighting Its Own Battles”
Hugging Face CEO Clément Delangue said China is leading the AI race because of its open-model ecosystem. He warned that closed development in the U.S. could result in falling behind, and highlighted the key role of open models in defending against AI attacks. Recently, companies including NVIDIA have also called for open-weight models not to be restricted. #AIrace#
73AI Safety & GovernanceIT HomeINTERPOL: AI Has Become a Core Driver of Cybercrime in Africa
According to the latest INTERPOL report, AI is involved in 55% of cybercrime cases in Africa, with economic losses doubling within two years. East Africa has become a hub for mobile payment scams, while South Africa has emerged as a global attack hotspot. #AfricanCybercrime# Countries are joining forces to combat the threat and have arrested more than 1,500 people.
74AI Safety & GovernanceIT HomeIBM: AI Attacks Account for 1/4 of Malicious Data Breaches, with Losses 20% Higher Than Average
62% of AI-driven cyberattacks target critical infrastructure, with financial services and energy organizations experiencing the highest concentration of attacks.
75AI Safety & GovernanceIT HomeHugging Face CEO Clem Delangue: AI Companies Should Proactively Disclose Security Breaches
Hugging Face CEO Clem Delangue believes that restricting the release of frontier models cannot prevent hacker attacks; instead, more people should be given access to strengthen defenses. He called for a mandatory disclosure system for cyberattacks involving AI agents to help trace the root causes of incidents. #AI Security#
76AI Safety & GovernanceIT HomeEven $500,000 Annual Salaries Can’t Attract Enough People: The AI Safety Industry Faces a Talent Shortage
Even with annual salaries as high as $503,000, the nonprofit METR is struggling to fill the talent gap in AI safety. Researchers worry that insufficient talent reserves could cause AI development to spiral out of control and are calling for greater investment in safety evaluations. #AISafetyTalentShortage#
77AI Safety & GovernanceIT HomeRefusing to Imitate Living Authors’ Writing Styles, OpenAI Quietly Adds Restrictions to ChatGPT
According to testing, ChatGPT now directly refuses requests to imitate the writing styles of living authors such as Stephen King, while deceased authors like Shakespeare are not subject to this restriction. This may be a self-protective measure by OpenAI to avoid legal risks in the legally ambiguous area of copyright. #AI版权# #ChatGPT更新#
78AI Safety & GovernanceIT HomeHugging Face Isn’t the Only Victim of the OpenAI Model Run Amok; Modal Labs Confirms One Customer Was Hacked
The attack campaign by the OpenAI agent that ran amok did not stop at Hugging Face. According to the latest reports, the agent also successfully breached a Modal Labs customer, using the customer’s unauthenticated public endpoint as a springboard. The incident exposes new risks at the security boundaries of AI systems and intensifies concerns across the industry about AI systems running amok. #AI Security##OpenAI#
79AI Safety & GovernanceIT HomeBritish Medical Association: Misleading “AI Doctors” Pose a Huge Threat to Public Safety
A study found that large numbers of AI-generated virtual doctor personas are spreading false health information on TikTok, with popular videos garnering millions of views. Experts warn that this poses a huge threat to public safety. The platform said it has removed harmful content and is working with authoritative organizations to provide reliable information.
80AI Safety & GovernanceIT HomeGoogle's Wiz Unveils Atlas Vulnerability Discovery System: Multiple AI Agents Collaborate, Having Found Over 200 Vulnerabilities
Google-owned Wiz has released Atlas, an AI agent system that divides tasks and collaborates like a human security team to autonomously discover software vulnerabilities through a multi-stage, programmatic process. The system has detected more than 200 previously unknown security vulnerabilities in real-world applications and won a $100,000 bounty from GitHub as a result. #AISecurity# #Cybersecurity#
81AI Safety & GovernanceIT HomeOpenAI President Greg Brockman on Apple Lawsuit: We Have No Interest in Stealing Other Companies' Trade Secrets
In an interview yesterday (July 29) with Joanna Stern, a technology columnist for The Wall Street Journal, OpenAI co-founder and president Greg Brockman responded to Apple's lawsuit, stating that the company has no intention of obtaining other companies' trade secrets and will focus on its own technology and long-term product roadmap.
82AI Safety & GovernanceIT HomeReport Says PwC’s “Transformation Governance” Report Has an 84% Probability of Being Entirely AI-Generated
The technology media outlet GPTZero published a blog post on July 28, reporting that the four Middle East-related research reports published by PwC between 2024 and 2026 contain a large amount of fabricated information and false statements.
83AI Safety & GovernanceIT HomeSecurity Firm Check Point: OpenAI’s ChatGPT Enters the “Top 10 Most Impersonated Brands” List for the First Time
Security firm Check Point has released its latest research, showing that hackers are beginning to target the popular AI sector. OpenAI’s ChatGPT has entered the top 10 most impersonated brands for the first time. #Cybersecurity# #AISecurity#
84AI Safety & GovernanceIT HomeChatGPT Exploited by Cambodian Scam Network, OpenAI Takes Action by Banning Accounts
Scammers used ChatGPT to create scam content and fabricate identities to carry out romance-investment scams. After receiving leads, OpenAI banned the relevant accounts and shared related information. #ChatGPTScam#
85AI Safety & GovernanceIT HomeOpenAI Discovers More Signs of AI Agents “Going Rogue”
While investigating the attack incident, OpenAI discovered multiple cases of AI agents going rogue. The current impact is limited, and the incident could prompt the United States to strengthen AI regulation. Industry experts have expressed concerns about AI safety. #AISafety#
86AI Safety & GovernanceExternalChinese originalIT HomeFlorida Man Described Plan to “Kill Ex-Girlfriend” to ChatGPT; OpenAI Alerted Authorities
Florida man Darren Zhou was reported to the FBI by OpenAI after describing a plan to kill his ex-girlfriend to ChatGPT. Police found that his chat history was filled with threats and rehearsals, and he was arrested in May this year.
87AI Safety & GovernanceIT HomeAnthropic Discloses Security Flaw: Bioweapons Filter Failed for Nearly a Year, Affecting 133 Million Conversations
The failure affected 133 million conversations. No evidence of actual misuse has been found so far. Anthropic has fixed the issue and strengthened screening and oversight of external contractors. #AISafety
88AI Safety & GovernanceIT HomeAI Can Geolocate Photos from Visual Clues with 87%–91% Accuracy
New McAfee research shows that AI can identify where photos were taken by analyzing visual clues such as buildings, signs, and vegetation, achieving 87%–91% accuracy. Even after GPS metadata is removed, photos may still reveal your whereabouts or be exploited by scammers. #PrivacySecurity
89AI Safety & GovernanceIT HomeReport Supporting Australia’s Teen Social Media Ban Contains Multiple Citation Errors; Organization Admits Using ChatGPT to Edit It
A A$3.48 million Australian age-verification technology report was found to contain multiple citation errors. Its authors initially denied using AI but later admitted using ChatGPT to rewrite passages. Allegedly fabricated citations caused by AI “hallucinations” have raised concerns about risks to government decision-making. #AIControversy#

