So the funny thing is, I've seen a lot of discussions about this, and most of them miss the point by fixating on the output alone. The real foundation for a tool like this is the input—the author's intent. It needs a ludicrously granular system for capturing tone and intent. Is this scene meant to be playful, desperate, melancholic, or absurdly over-the-top? A simple dropdown menu for 'mood' won't cut it. It needs to be a collaborative filter that asks, 'What's the emotional trajectory here? What's the power dynamic at this exact moment?' Without that, you just get mechanically generated physical descriptions that feel completely disconnected from character.
It also desperately needs a 'consistency engine' that's less about remembering eye color and more about tracking emotional and physical continuity. If Character A is shy and hesitant at the start of an encounter, the AI needs to understand how that shyness might dissolve or transform, not just forget it and make them suddenly dominant three paragraphs later. That break in character ruins immersion faster than anything. The tech should help you maintain the internal logic of the scene, which is what actually makes the writing resonate, or not.
Finally, it must have robust, invisible safety features for the creator. Not clumsy censorship, but the ability to set personal boundaries for content you don't want to accidentally generate or see in suggestions. And everything, from the training data upward, should be ethically sourced with clear creator opt-in. Otherwise, you're building on a pretty grim foundation.