Close Menu
arabianstartup.comarabianstartup.com
    What's Hot

    Here’s the latest.

    October 13, 2025

    Why Now? The Lost Chances to Reach a Hostage Deal, and a Cease-Fire, Months Ago

    October 12, 2025

    The ZoraSafe app wants to protect older people online and will present at TechCrunch Disrupt 2025 

    October 12, 2025
    Facebook X (Twitter) Instagram
    Facebook X (Twitter) Instagram
    arabianstartup.comarabianstartup.com
    Subscribe
    • Home
    • Insights
    • Business
    • Feature
    • Market Trend
    • Startups
    arabianstartup.comarabianstartup.com
    Home » Anthropic says some Claude models can now end ‘harmful or abusive’ conversations 
    Startups

    Anthropic says some Claude models can now end ‘harmful or abusive’ conversations 

    Arabian Media staffBy Arabian Media staffAugust 16, 2025No Comments2 Mins Read
    Facebook Twitter LinkedIn Telegram Pinterest Tumblr Reddit WhatsApp Email
    Share
    Facebook Twitter LinkedIn Pinterest Email


    Anthropic has announced new capabilities that will allow some of its newest, largest models to end conversations in what the company describes as “rare, extreme cases of persistently harmful or abusive user interactions.” Strikingly, Anthropic says it’s doing this not to protect the human user, but rather the AI model itself.

    To be clear, the company isn’t claiming that its Claude AI models are sentient or can be harmed by their conversations with users. In its own words, Anthropic remains “highly uncertain about the potential moral status of Claude and other LLMs, now or in the future.”

    However, its announcement points to a recent program created to study what it calls “model welfare” and says Anthropic is essentially taking a just-in-case approach, “working to identify and implement low-cost interventions to mitigate risks to model welfare, in case such welfare is possible.”

    This latest change is currently limited to Claude Opus 4 and 4.1. And again, it’s only supposed to happen in “extreme edge cases,” such as “requests from users for sexual content involving minors and attempts to solicit information that would enable large-scale violence or acts of terror.”

    While those types of requests could potentially create legal or publicity problems for Anthropic itself (witness recent reporting around how ChatGPT can potentially reinforce or contribute to its users’ delusional thinking), the company says that in pre-deployment testing, Claude Opus 4 showed a “strong preference against” responding to these requests and a “pattern of apparent distress” when it did so.

    As for these new conversation-ending capabilities, the company says, “In all cases, Claude is only to use its conversation-ending ability as a last resort when multiple attempts at redirection have failed and hope of a productive interaction has been exhausted, or when a user explicitly asks Claude to end a chat.”

    Anthropic also says Claude has been “directed not to use this ability in cases where users might be at imminent risk of harming themselves or others.”

    Techcrunch event

    San Francisco
    |
    October 27-29, 2025

    When Claude does end a conversation, Anthropic says users will still be able to start new conversations from the same account, and to create new branches of the troublesome conversation by editing their responses.

    “We’re treating this feature as an ongoing experiment and will continue refining our approach,” the company says.



    Source link

    Share. Facebook Twitter Pinterest LinkedIn Tumblr Telegram Email
    Previous ArticleIsrael Gears Up for Nationwide Strike to Support Hostages
    Next Article U.S. Pauses Visitor Visas for Gazans After Right-Wing Outcry
    Arabian Media staff
    • Website

    Related Posts

    The ZoraSafe app wants to protect older people online and will present at TechCrunch Disrupt 2025 

    October 12, 2025

    Nvidia’s AI empire: A look at its top startup investments

    October 12, 2025

    Dating app Cerca will show how Gen Z really dates at TechCrunch Disrupt 2025

    October 12, 2025
    Add A Comment
    Leave A Reply Cancel Reply

    Top Posts

    10 Trends From Year 2020 That Predict Business Apps Popularity

    January 20, 2021

    Shipping Lines Continue to Increase Fees, Firms Face More Difficulties

    January 15, 2021

    Qatar Airways Helps Bring Tens of Thousands of Seafarers

    January 15, 2021

    Subscribe to Updates

    Unlock the latest trends, insights, and expert advice in the world of startups and entrepreneurship with our exclusive newsletter.

    Welcome to Arabian Startup, your ultimate source for the latest trends, insights, and success stories in the world of startups and entrepreneurship.

    Facebook X (Twitter) Instagram Pinterest YouTube
    Top Insights

    Top UK Stocks to Watch: Capita Shares Rise as it Unveils

    January 15, 2021
    8.5

    Digital Euro Might Suck Away 8% of Banks’ Deposits

    January 12, 2021

    Oil Gains on OPEC Outlook That U.S. Growth Will Slow

    January 11, 2021
    Get Informed

    Subscribe to Updates

    Unlock the latest trends, insights, and expert advice in the world of startups and entrepreneurship with our exclusive newsletter.

    @2025 copyright by Arabian Media Group
    • Home
    • About Us

    Type above and press Enter to search. Press Esc to cancel.