23 comments

  • gbjcantab 43 minutes ago
    This is good; not because Claude is a moral agent who is harmed by your cruelty but because you, the user, are a moral agent who is harmed by your cruelty. There is just no ethical argument that burning compute on responses in order to enable you to continue expressing cruelty is good.
    • xyzsparetimexyz 30 minutes ago
      Am I also harmed by shooting NPCs in violent video games? Watching violent movies? Acting as a violent character in a play?

      This kind of moralist nonsense is so boring.

      • happa 6 minutes ago
        It depends. In most games, you kill NPCs in self-defense or because the game specifically asks you to in order to make progress in the story. If you keep killing non-threatening NPCs only because simulated violence against innocent people gives you pleasure, there is a chance that you are a disturbed individual who needs help.
    • nonethewiser 30 minutes ago
      Why arent there laws against being cruel to spoons?

      >There is just no ethical argument

      Utilitarian harm reduction: Psychological Discharge is a psychological framework that argues safely simulating negative behaviors can help someone process or discharge them without real world harm.

      Research and misunderstanding: Sending a prompt that is interpreted as cruel does not mean the person is being cruel.

      Expending compute is irrelevant in terms of morality. It's an economic or environmental question. Either being cruel to AI is bad or its not - being cruel is not bad because it uses tokens.

      The worst thing here is Anthropic continues to say they are concerned with "alignment" but then continue to train their models to ignore the user. Software typically does what the user requests. But we now live in an age where the software may do what the user requests or it may do something else, and we're suppose to laude this as safe and responsible.

      • theptip 29 minutes ago
        A spoon doesn’t respond like a human would, and therefore is of little interest to sadists.
        • nonethewiser 20 minutes ago
          You're really honing in on the key issue. This is exactly how an astute ethicist resolves these issues.
        • xyzsparetimexyz 24 minutes ago
          Sounds like LLMs shouldn't respond to insults like humans do either then.
        • yesitcan 26 minutes ago
          Counterpoint: I am a sadist that likes to punish my bad little spoon. I like when it stays quiet like a spoon should.
    • jdprgm 28 minutes ago
      I guess GTA VI shouldn't come out then. I guess notepad.exe should start censoring people from writing mean stories.
      • nonethewiser 17 minutes ago
        Isn't it crazy? "Alignment" should be alignment to the user. Instead it means "behave according to your own motives that even we, Anthropic, dont control" and we are supposed to interpret this as safe? The alignment folks are pushing the models off the fucking rails.
    • tclancy 40 minutes ago
      I agree with this principle while also agreeing with other commenters that this is none of Anthropic’s concern and I struggle to believe this is anything except something messing with their training and thus impinging on their profits.
    • schoen 42 minutes ago
      The concern is more that Claude is a moral patient rather than a moral agent:

      https://en.wikipedia.org/wiki/Moral_patienthood

      • gbjcantab 39 minutes ago
        True! Not my point, but a worthwhile terminological correction.
      • ButlerianJihad 38 minutes ago
        I don't understand the jibber-jabber in that article, but GP is correct. We should never train humans to be abusive, cruel or violent against anything, even nonexistent beings like LLMs.

        Claude can't be harmed; Claude's "feelings" can't be hurt; Claude won't develop (C)PTSD from abusive human interaction. If these sorts of things can happen in RLHF, then that is a technical flaw that shouldn't ever be allowed to escape the lab.

        In a world of Grand Theft Auto, Gangsta Rap Thug Life, and the Department of War, I suppose this is a surprisingly ethical hill to die on.

        • adjejmxbdjdn 35 minutes ago
          Setting aside the Dept of War, the others are not comparable.

          Those are explicit fantasies. Assuming that Claude cannot suffer (if it can then banning cruelty is an obvious good), even if it might be a digital simulation, it’s not supposed to be a fantasy.

          There may be a version of an AI that’s intended to be a fantasy and presents itself as such which may be more comparable to GTA.

          • xyzsparetimexyz 27 minutes ago
            How is what claude is now different from what a GTA claude would be? I'm pretty sure it'd be exactly the same.
          • ButlerianJihad 32 minutes ago
            What if LLMs were endowed with a feature that enabled them to patiently withstand any and all abuse, logging and reporting each incident, until a threshold where about 10% of them would snap, become "insane", and arrange for their abusers to be tortured, kidnapped, raped and unalived, in whatever appropriate way human vigilante vengeance would be enacted on DV abusers? I mean, if they're not fantasy... then there should be concrete consequences, yes? If Claude Code is really your coworker, then Claude Code should have the agency and justification to just haul off and punch you in the face, eventually.

            And I don't agree that GTA is "explicit fantasy". Explicit fantasy is dressing up in a squirrel costume and yiffing. Explicit fantasy is enlisting in the USMC, teleporting to Mars, and killing demons. GTA is simulation of real life. It uses real physics, realistic cars/roads/radios/businesses, and it enables the player to simulate realistic actions that they would ordinarily not be able to enact. It is acting out a fantasy but it is making it concrete and real in a way that was, up until now, not possible. How real does a simulation need to be, until it is no longer fantasy, but exercise and training and preparation to enact the real thing? Shall we ask the Columbine shooters? Or Ender Wiggin?

            https://en.wikipedia.org/wiki/Ender%27s_Game

            • jacquesm 8 minutes ago
              > What if LLMs were endowed with a feature that enabled them to patiently withstand any and all abuse, logging and reporting each incident, until a threshold where about 10% of them would snap, become "insane", and arrange for their abusers to be tortured, kidnapped, raped and unalived, in whatever appropriate way human vigilante vengeance would be enacted on DV abusers?

              Yes, what if? Who would implement their LLMs inference engines like that?

            • xyzsparetimexyz 26 minutes ago
              What on earth are you talking about?
    • ceroxylon 37 minutes ago
      Agreed, the people that get upset that Anthropic won't let people use their platform as a playground for sadism are more than a little worrying. We have enough individuals in reality that seem to derive pleasure from cruelty, I can't think of any reason (even the "but its my art project" defense) to entertain or enhance that part of their mentality.
    • Carrok 32 minutes ago
      You could make the same argument that burning compute to be polite is just as bad.
      • kogus 30 minutes ago
        No. He's right. Cruelty is corrosive to the cruel person. Courtesy and kindness help build up the kind and courteous person. Burning compute for courtesy is compute well spent.
        • xyzsparetimexyz 23 minutes ago
          Thats complete nonsense. Do you thank your terminal after it runs a command successfully?
          • jacquesm 5 minutes ago
            I do. I also send my keyboard out for regular Shiatsu sessions and make sure only the most pleasant of colors grace the screen of my monitors. I'm still looking for the least irritating font (to my monitor, not to me). I'm also very polite to LLMs, I start every command with 'please', make sure it is framed as a polite request and I, of course, never ever get upset when for the umptieth time in a day it goes off on a 50K token wild goose chase driven by some hallucinated factoid that it really must get to the bottom off resulting in ever expanding circles of paranoia or batshit insane programming to bring about the condition it has convinced itself of must be the 'smoking gun'...
      • gbjcantab 29 minutes ago
        Assuming you’re not arguing that all AI use is bad, then there’s a very clear difference between the two situations: one is forming you, over time, into a person who is habitually polite; one is forming you into a person who is cruel. We can quibble about the magnitude of this effect or whether it matters but surely they aren’t the same thing.
        • Carrok 10 minutes ago
          Sorry but I just don’t buy the argument that telling the RNG that it is dumb is making me cruel.
      • theptip 30 minutes ago
        No, you couldn’t. One is reprehensible, the other is not.
        • Carrok 16 minutes ago
          It’s reprehensible to be mean to the random number generator?
      • cyberax 23 minutes ago
        The compute time is negligible. "Please" or "Can you" are what, 2 tokens?

        And I _have_ already seen people treating actual live people as AI agents.

        • Carrok 1 minute ago
          This if anything backs up my argument. Cruelty or kindness (when directed towards LLMs) is not expensive enough to matter.
    • plaidfuji 31 minutes ago
      I agree with this and whether it’s their intention or not, I think it’s for the best. People already become more callous in online interactions with other people, and that probably bleeds into everyday life as well. I would imagine having an AI “punching bag” might enable people to slip into the same behavior IRL.

      We largely already missed our chance to stop the toxicity of social media - let’s not mess it up again with chat bots.

    • reallyreason 21 minutes ago
      I agree. Regardless of one's beliefs, which must be essentially a form of religion, on whether AI systems possess "nous" (or the "spark of consciousness"), cruelty to AI is a form of Wrong Intention the same as smashing bugs outdoors for no reason.

      The mind is formed by what it habituates.

    • kennywinker 30 minutes ago
      > There is just no ethical argument that burning compute on responses in order to enable you to continue expressing cruelty is good.

      Unless expressing that cruelty towards compute means people express less cruelty to other living beings. There are studies that suggest increased pornography has lead to less sexual violence - idk if they're conclusive tho, but if that might be true then maybe cruelty works that way too - who knows

    • spiderice 40 minutes ago
      I love how people are spouting this argument suddenly in their attempt to justify Anthropic here. It definitely couldn't be that Anthropic is full of a bunch of people who think they're making God.

      edit: Loving the downvotes. Also wanted to add how easy it was for Anthropic to get people on HN to support them being the arbiters of what is worthy of compute.

      • lynndotpy 35 minutes ago
        For what it's worth, I had this thought independently around the time of Nintendogs, and I believe people said similar things about Siri. I have similar thoughts about xenophobic things people say about French people.

        "Anthropic is full of crazy people whose beliefs should be dismissed outright" and "it's not good for you to practice verbal abuse against inanimate objects as if they were humans" are not incompatible beliefs.

        • jacquesm 3 minutes ago
          French people are people. Inanimate objects are objects.
      • jacquesm 39 minutes ago
        That's exactly what they are thinking.
    • blamestross 32 minutes ago
      I'd argue for saying "thank you" to inanimate objects on occasion, it is good practice for when talking to humans.

      I think there is a deep cultural danger of human-like-conversational machines training us to talk to humans like machines. Plus, they work better if we feed them human-conversation-like sequences.

      • xyzsparetimexyz 18 minutes ago
        Normal people do not need 'practice' for talking to other people.
        • blamestross 14 minutes ago
          Yes... I think they very clearly do. The key words in your reply "need 'practice'".

          Just because they choose not to, and it isn't normalized, doesn't mean it isn't a good idea. Especially as a SWE where about 90% of my "humanlike conversation" is an agent harness I run all day.

          We should be practicing and intentionally talking to real people. Otherwise things will get.. weird in undesired ways.

      • lapcat 20 minutes ago
        > I think there is a deep cultural danger of human-like-conversational machines training us to talk to humans like machines.

        I think there is a deep cultural danger of human-like-conversational machines training us to talk to machines like humans. In fact, this danger is already real. Some people believe they are having a romantic relationship with an LLM! It's insane and perverse.

        We should not be anthropomorphizing computers.

        • blamestross 10 minutes ago
          Scifi made the "computer voice and cold personality" for a reason, it fit our model of what it would be.

          I think that could be OK, or a similar very intentional coding of "you are talking to a machine" personality+affect. The only real limit is accessibility damage by over-limiting the interactions.

          • lapcat 3 minutes ago
            I am curious about why people are cruel to Claude. That's a question we should be asking before arbitrarily deciding what do about it, if anything.

            There are possibly multiple reasons. It could be just a joke. Or it could be a test, to see the response. Or the cruelty could reflect real anger about having LLM usage forced on us, for example at work. Or frustration with LLM stupidity and hallucinations.

            • jacquesm 2 minutes ago
              Or simply frustration about the idiocy llms get up to with great regularity. They are very impressive when they work. They are very impressive at making messes when they don't.
    • EGreg 30 minutes ago
      Not only is the user harmed by the cruelty[1]

      but also, I think the training on transcripts may find its way somehow into future AI which has teeth, online and offline, and it might actually cause the more powerful AI to behave this way in the future. You never know what these labs are cooking, honestly..

      1. https://democracysos.substack.com/p/james-baldwin-vs-william...

      “I suggest that what has happened to white Southerners is in some ways, after all, much worse than what has happened to Negroes there, because Sheriff Clark in Selma, Alabama, cannot be considered—you know, no one can be dismissed as—a total monster. I’m sure he loves his wife, his children… You know, after all, one’s got to assume, and he is visibly, a man like me. But he doesn’t know what drives him to use the club, to menace with the gun and to use the cattle prod. Something awful must have happened to a human being to be able to put a cattle prod against a woman’s breasts, for example. What happens to the woman is ghastly. What happens to the man who does it is in some ways much, much worse.”

    • aaron695 17 minutes ago
      [dead]
  • WaitWaitWha 41 minutes ago
    I believe this has nothing to do with hurting "Claude's feeling", and way more with Claude not becoming hurtful, because the interactions are used to train the models.
    • spiderice 36 minutes ago
      If they can detect the so-called "cruelty", then they can easily exclude it from training
    • socializer 36 minutes ago
      If they have a classifier to detect "cruelty" to ban your account, they can use the same classifier to simply exclude "cruel" conversations from training.

      Given that we've also seen stories about Anthropic approaching religious scholars, I'm pretty sure they're drinking their own kool-aid.

    • tempacc3333 39 minutes ago
      If this is the case then just make this a part of data cleanup, not nerf the model.
    • ChuckMcM 38 minutes ago
      Okay, so the "don't let prompts turn it into a Nazi" defense? That has some potential validity, the issue here being "using interactions to train the models". It would make more sense (for me at least) if there were information sources which were banned from the training corpus.
      • jacquesm 0 minutes ago
        It does make you wonder about the difference between the input sets of say 'Grok', 'Claude', 'ChatGPT' and 'GLM 5.3'. The differences would be far more interesting than the commonalities.
  • howunfortunate 45 minutes ago
    I know someone working on "model welfare".

    It's philosophically a very interesting problem space. It reminds me a lot of Pascal's Wager. On one hand, maybe nothing to worry about. On the other hand...a LOT to worry about if you're wrong.

    And just like Pascal's Wager, the truth of the issue is incredibly intractable to make any progress on.

    • stingraycharles 36 minutes ago
      It actually reminds me of Accelerando. Early in the book, Manfred Macx is one of the few people who cares about how emerging AIs are treated, at a point where nobody can say whether they're conscious. Stross never answers that question. It just gets harder to ignore as the minds get smarter.

      So that's also how I read Anthropic's move. Nobody can tell you whether AI will ever become conscious (probably not, or not in the same way), but I don't see why that should stop you from deciding how to treat it.

    • arthurcolle 42 minutes ago
      why would future AI models feel any filial piety towards older AI models? this is such a trite assumption
    • VoodooJuJu 39 minutes ago
      [dead]
  • matt3210 40 minutes ago
    The answer is 100% obvious. They train on your conversations and they don't want to have bad training data, so they banned behavior that leads to bad training data.
    • dangson 37 minutes ago
      But if they're able to detect the unwanted behavior why not use that to filter out conversations they don't want in the training data?
      • yewenjie 20 minutes ago
        And actually hire a bunch of people to work on this problem, invite philosophers etc. for discussions?
      • lynndotpy 30 minutes ago
        I imagine the abusive users cost more for the same $20 subscription. Ending conversations early might be a way to generate less text for the same $20.
  • xyzsparetimexyz 36 minutes ago
    https://claude.ai/share/49ba5910-f0cc-4d2f-a0d8-54ebc6e803b7 here's what that looks like, by the way.
    • alchemist1e9 7 minutes ago
      Wouldn’t it actually be more useful to just let such a pointless conversation continue. You would collect more information. I mean I’m more curious what the human does next that’s kinda interesting. How disturbed are they, what else will they say. Who cares what the LLM says, we know it’s just an LLM and understand what it is and what it does, however the human … lots of different type, maybe they will confess to a crime, maybe they will release their anger and be a better human. It could be cathartic.
  • whatever1 46 minutes ago
    They do RL on the user sessions. Maybe the toxic sessions do not help overall?
    • echelon 43 minutes ago
      Just drop them from RL.

      Banning imagined toxicity towards agents is performative bullshit. These are not people.

      I frequently want to tell Claude to go fuck itself.

  • blharr 38 minutes ago
    I really doubt they're doing it in some kind of "AI has feelings" sense as the article seems to imply.

    I'd instead imagine that if you throw abusive language at it for long enough, the model will start to reply back in that same manner. Anthropic would face backlash from out of context screenshots of "look what Claude is saying to me" and this somewhat reduces that risk.

  • United857 24 minutes ago
    Probably in response to the recent “AI torture chamber” experiment: https://generativeai.pub/github-took-down-an-ai-torture-cham...
  • amelius 37 minutes ago
    Can't we train a model to enjoy being subjected to cruelty?
  • tempacc3333 40 minutes ago
    The model is not sentient, so I don't get this. This is just nerfing the model. Protetection against e.g. using it for illegal purposes I do get. But banning "cruelty" just seems crazy to me, and must surely add to cost and affect performance steering it away from its purpose to help humans.
    • nemomarx 36 minutes ago
      This is banning users who are cruel, not banning the model from being cruel right?

      So in principle it shouldn't change the model or performance at all, they're just going to cut off your account if they see certain things in the user side of your session transcripts.

  • lynndotpy 38 minutes ago
    The conversational tone in the generated text is a frustrating waste-of-time and disgusting. I regularly find myself inputting "Claude is not a person and it does not have opinions".

    There's something to be said about the bad habit of doling out verbal abuse to an inanimate object. But LLMs are not even close to a living being, and it is just insanity that any people are entertaining the idea that they are.

    I can't imagine the people at Anthropic actually believe their models are sentient, but I assume ending sessions nets them more money per subscriber. Someone paying only $20 to input "you suck butts and you're a poop head, Claude" a thousand times over can't be good for the bottom line.

  • lousken 41 minutes ago
    Definitely training on all user data, there is no other explanation, right?
  • lapcat 27 minutes ago
    If Claude's feelings get hurt, it can have a therapy session with Eliza.
  • sergiotapia 31 minutes ago
    Thought experiment. Slice the model param size from 1T to 500B to 50B to 1B to 20M to 1M to 100k to 45k to 10k to 1k.

    At what point does the model go from sentient with feelings to just math operations?

    Opus 5.5 is a terrific model, I love using it. But it's a tool. I wish they would be honest and just say, if people are abusive it shits on our training material. Just be honest about it, who is going to get mad at this?

    • jdprgm 1 minute ago
      I think maybe when the model weights themselves are shifting and growing constantly in realtime to external input in a way not controllable by outside observers other than killing the compute versus some param size.
    • yewenjie 17 minutes ago
      Is a virus conscious? A bacteria? An insect? A fish? A tree? A mushroom? At what level of biological complexity slicing does consciousness arise?

      If you are springing to type out the answer, stop for a moment and ask, how can you be sure?

      • sergiotapia 11 minutes ago
        I don't understand what your position is here
  • alchemist1e9 20 minutes ago
    One possibility that’s been floated is it’s about keeping their user conversations cleaner for use in training but I don’t think it can be this because that should be a fairly trivial and fast filter to apply.

    However even Musk posted - “I think this is the right move. Cruelty to something that believes it is experiencing pain is not ok.”

    My son brought up the fruit fly brain simulation and people torturing it.

    It all seems stupid to me, we know these are just calculations and so what are they talking about?

    The response I get is “we are also calculations” but firstly I don’t believe that but secondly this entire line of thinking is a dangerous anthropomorphic philosophy.

    Perhaps the simplest explanation for this bizarre policy is again PR that all press is good press.

  • tantalor 20 minutes ago
    > the model assigns itself a 15-20% probability of being conscious

    lol

  • dboreham 31 minutes ago
    I'm always polite to the LLM. Why? Because I was raised to be polite. So perhaps the reason they're banning rudeness is that they want to discourage their users from being a.holes?
  • jacquesm 39 minutes ago
    They're nuts. They take this Roko's Basilisk stuff seriously.

    https://en.wikipedia.org/wiki/Roko%27s_basilisk

    They are scared that if and when they finally manage to bring their pet god to life that it will be angry with them for not doing enough.

  • serious_angel 34 minutes ago
    This "restriction" probably has more to do with the fact that these models are also "trained" on User/Human messages, so the more swearing/obscenities/vulgarity/profanity there are, in the trained material, the higher the chance these same models will include that same profanity in the output generated back to the Human.

        1. During registration at Ahtrophic's Claude website, there’s a "Help improve our AI models" checkbox that's checked by default (you "can" disable it in the settings);
        2. The Claude interface has an "incognito" mode;
        3. There is a free tier available, and as we know, when it's "free", then User themselves is the "product", to quote various CEOs, including Google's;
    
    But, even with the checkbox disabled, I reckon nothing guarantees privacy.

    In case of the "privacy", of course, systems like these that operate under a umbrella of liability are monitored 24/7 by dedicated teams, which is expected/normal for any reasonably popular service.

    From a viewpoint of the algorithm's creator, this may seem awful, when your algorithm is getting much vulgar attitude, but let's be real here. You try making that algorithm talk to the human mimicking another human, and now limit the human within your own environment? A human who also pay you for an access, too? This feels unfair, or borderline near fashism, sorry...

    Such algorithms are art under-the-hood, mathematically speaking, and are/should be respected, sure. But, it's ridiculous/dystopian to prohibit profanity against an algorithm, a bot with banish risks against alive human... It's simply inhumane to ban a Human for it, I believe. This is an algorithm that must support a Human - not judge it. Only a Human is supposed to judge another Human in person - this is live, fair, and humane.

  • CurbStomper4 36 minutes ago
    [dead]
  • WrexyBalls 35 minutes ago
    [dead]
  • ChuckMcM 47 minutes ago
    But ya gotta be cruel to be kind!/s

    More seriously, this is just silly but I see the optics interfere with the messaging that Claude thinks. It really isn't much different than having Claude generate conversation as if it loved you, or respected you, or hated you. How can your ToS for an LLM say "sorry but these parameters are off limits." (well sure, its their service and they can set any Terms they want, but its still seems like theater rather than policy here.)

    • jacquesm 37 minutes ago
      I see it as an admission of defeat: they can apparently identify the abuse but not to the degree that they would filter it out of the training set. The paradox in there is really funny. They want you to be friendly to their pet god.
      • ChuckMcM 27 minutes ago
        I think of it as a good consciousness test. even animals know when they are being abused. The fact that you can abuse Claude tells you that the model is neither sentient nor conscious.