StoryScope’s paper scores 61,608 stories: 10,272 human Books3 texts and up to five LLM mirrors per prompt. Models refused 24 generations. Narrative features alone reach 93.2% macro-F1 for human vs. AI. On 278 Gemini stories, LAMP rewrote seven classes of surface artifacts (cliché, redundant exposition, purple prose) with Gemini as the rewriter. The narrative detector moved from 95.5% to 93.9% macro-F1.

The 30-feature core still reaches 84.8% on the main binary task. After those surface rewrites, the paper still points at theme over-explanation, single-track plots, and linear time. Human stories more often leave the moral implicit. They also break chronology.

I would treat a banned-phrase list as the wrong lever. Score those three, and say whether the text was rewritten for style.

source ↗

← all notes