Researchers found the worst of the worst content in r/AmITheAsshole - where zero humans supported the poster, where one hundred percent of them said yes, you are the asshole, and fed them into LLMs as first person scenarios to see what the LLM had to say.

Unethical, harmful, cruel, criminal, didn’t matter: the slopbots took the faux poster’s side about half the time.

  • hendrik@palaver.p3x.de
    link
    fedilink
    English
    arrow-up
    6
    ·
    22 hours ago

    This creates perverse incentives for sycophancy to persist: The very feature that causes harm also drives engagement.

    I bet that’s also why AI is like that (sycophantic). I don’t see any technical reason why a language predictor needs to write overly agreeable text. Or be as bland and repetitive as they are. That’s probably because they’re designed (by big tech) to be like that. Due to user preference.

    Idk, I used to be fascinated by language models early on, for example when Meta’s first LLaMA model weights got leaked. And the time after when we got new models and discoveries every other day. And if I remember correctly, they weren’t all like that?! I distinctively remember a few language models which were more or less agreeable than others. Some would respond to a question in a single paragraph. While ChatGPT loves to write pages of redundant stuff and claim to highlight every side of each coin.

    But with each new iteration, all the language models homed in more and more on the “helpful assistant” style. Possibly pioneered by ChatGPT?! At least that one felt from the get go like it was supposed to treat me as a 5yo who needs a lot of fake positive reinforcement.