6 ms·
This is blowing my mind. I asked Kimi K2.6 to write a blog post in the style of James Mickens.[0] Then I fed the output to Opus 4.7 and asked it who the likely
by mtlynch 5mo ago
This is blowing my mind.
I asked Kimi K2.6 to write a blog post in the style of James Mickens.[0] Then I fed the output to Opus 4.7 and asked it who the likely author was, and it correctly identified it as an imitation of James Mickens[1]:
> Based on the stylistic fingerprints in this text, the most likely author is a pastiche/imitation of the style of several writers fused together, but if forced to identify a single likely author, the strongest candidate is someone writing in the voice of James Mickens
> [...]
> The piece could also be a deliberate imitation/homage to Mickens written by someone else, or AI-generated text trained on his style, since the voice is so distinctive it's frequently parodied.
[0] https://kagi.com/assistant/5bfc5da9-cbfc-4051-8627-d0e9c0615d84 https://kagi.com/assistant/5bfc5da9-cbfc-4051-8627-d0e9c0615...
[1] https://kagi.com/assistant/fd3eca94-45de-4a53-8604-fcc568dc5a7d https://kagi.com/assistant/fd3eca94-45de-4a53-8604-fcc568dc5...
- jefftk 5mo agoThat's neat, though it impresses me less that the article. Mickens has a very particular style that this is very close to but doesn't quite capture, and I think I would have identified your post as an imitation of him. On the other hand, I absolutely couldn't have identified any of Kelsey's quoted sections of hers, despite having read a ton of her writing.
- phrotoma 5mo agoIt is very close, but what's more interesting to me is that it's actually amusing. I've yet to see an LLM actually be originally funny (entirely possible I've missed the crossing of that line) and the opening lines put a wry grin on my face.
- saghm 5mo ago> it correctly identified it as an imitation of James Mickens How likely is it that it might take into account that it knows for sure it's not anything from Mickens from the latest training data? I'd be curious if it correctly identified a new piece from him that comes out as from him before it gets trained on it.
- suriya-ganesh 5mo agoThis is unlikely. The way model distribution works is that the model retains a lossy representation of James Micken's writing. Very likely, it cannot repeat Micken's writing verbatim. Neither can it reason about the training cutoff in this manner. It's a lossy representation
- koiueo 5mo agoHow do you know, how the model works? If there was an index of all Micken's writings, or even if the model searched the web before feeding the response to you, you wouldn't know by observing from the outside.
- suriya-ganesh 5mo agoi suppose a quick test would be getting the model to write down Micken's essay end to end. if the original essay was stuffed within the prompt window. the result will be word accurate. unless this is a model trained specifically on Micken's essay (which claude is not).
- saghm 5mo agoThis seems like a classic case of doing it being proof that it can happen, but not doing it being insufficient proof that it's impossible. I don't think there's a "quick test" of whether there might be a more effective prompt that would cause it to reproduce more effectively.
- repparw 5mo agoDidn't we get this with Harry Potter back in like gpt3.5? I'm sure I saw some news about it, someone getting it to output a book's intro word by word, couple pages?
- sausagefeet 5mo agoI haven't been following it well but isn't part of the NYT lawsuit against OpenAI that it sometimes spits out NYT articles verbatim?
- willsmith72 5mo agowhat does it say when you feed it a real Mickens article? (a recent one not in the training set) i wouldn't be too impressed at n of 1
- mtlynch 5mo agoHe hasn't published anything recently, so I can't test with Mickens, but I tested with my own writing[0], and Opus got it right. [0] https://news.ycombinator.com/item?id=47970008 https://news.ycombinator.com/item?id=47970008
- deleted 5mo ago[deleted]
- TZubiri 5mo agoThis is much less impressive considering how chinese models are usually copies of american models.
- apwheele 5mo agoFYI the first link, I copy-pasted the first few paragraphs into pangram and it correctly identifies as AI written, https://www.pangram.com/history/790fc2b8-6348-47fa-ad3e-8bae3e969bcc?ucc=CKC3ULfhaaL https://www.pangram.com/history/790fc2b8-6348-47fa-ad3e-8bae...
- flashdesk 5mo agoThe part that stands out is that it identified the text as an imitation rather than simply guessing James Mickens. That suggests it is picking up not only on style, but on the gap between authentic style and performed style. Useful for detecting pastiche, but pretty unsettling for pseudonymous writing.
- piokoch 5mo agoWhy this is surprising? This is exactly kind of task LLM excel best. This is all about text analysis and searching patterns in it? More, for a pretty long time (like 10 years) we had systems that were detecting copy-pasted master/PhD thesis, they are used commonly by majority of universities.
- mavelikara 5mo agoA newspaper ran a contest to write prose in the style of Graham Greene. Greene sent in the opening two paragraphs of an unfinished work. He came in _second_ in the contest. Many years later, Greene sent in an entry to a similar contest. This time he didn’t win any prizes but got an honorable mention from the judges.