How to Ask the Wrong Question About AI Creativity

If I see another article speculating about the “creative potential” of LLMs, I think I’m going to give up the internet and make good on my promise to become a shrimp biologist.
I have a theory about how we got here, and it has to do with a historical misunderstanding.
There’s a conventional story that goes like this: Ada Lovelace, mathematician and writer, is widely regarded as the first computer programmer for her work on Charles Babbage’s theoretical “analytical engine.” The specific nature of her contributions are less popularly known, but if you’ve encountered her words at all, it’s likely this:
The Analytical Engine has no pretensions whatever to originate anything. It can do whatever we know how to order it to perform. It can follow analysis; but it has no power of anticipating any analytical relations or truths.
This excerpt is famously quoted by Alan Turing in his paper "Computing Machinery and Intelligence" in which he considers the question: “Can machines think?” He disagrees with Lovelace and goes on to propose his own standard for whether computers can think in his famous “Imitation Game.” In it, a judge in a blind test is asked to assess whether they are communicating with a computer or human. The ability to fool the judge, he states, might be taken as evidence of human cognitive capacity, contrary to the “Lovelace objection” to artificial intelligence.
[An interesting aside, since gender is at the periphery of this whole conversation: the game Turing proposes is based off of a party game in which a judge is supposed to determine whether they are engaging with the text of a man or woman behind a closed door. And an aside to my aside: many have written about the gendering of social robots (which follows typical stereotypes — troubling when women are so often under-represented in the development of technology), especially of personal assistants like Siri, which are disproportionately portrayed as women.]
The Lovelace objection was one among eight other objections Turing set out to debunk. An attentive reader can discern what’s happening here: Turing substitutes “surprise” for “origination” – a considerably lower criterion to debunk than Lovelace originally set, as “success” is then gauged by the judge’s reaction rather than what a machine is actually doing. His wasn’t a careful reading, nor was it intended to be a dialogue. Nevertheless, because of this citation their names remain twisted together in present-day discussions of computational creativity.
We took Turing’s proposition and flew with it. The so-called Turing Test has been invoked countless times as a sort of pseudo-benchmark in AI development, especially for language-based AI models. For years, the annual Loebner Prize was granted to computer programs that could trick judges into thinking they were human. “Passing the Turing Test” became a shorthand way of communicating the uncanny, conscious-seeming outputs of popular language-based models, from Joseph Weizenbaum’s 1966 ELIZA therapeutic chatbot, to Google’s LaMDA (made famous when in 2022 engineer Blake Lemoine claimed it achieved sentience) to ChatGPT. Even CAPTCHA challenges function as a sort of Turing Test, asking a human to interpret something that a robot theoretically could not.
It makes sense why this interpretation of the Turing Test captured popular imagination. It’s playful and undeniably elegant. Anyone can comprehend and deploy it. It’s also complementary to our deep history of speculating about artificial consciousness, which so often takes the shape of our own.
By now, for many, the shortcomings of the Turing standard are evident. If the Loebner Prize weren’t already discontinued after 2019, the recent generations of LLMs would render it entirely obsolete, and despite well-publicized cases of AI psychosis, the vast majority of people agree that ChatGPT is not conscious. Computer scientists are quick to point out that the Turing Test was never taken as a serious benchmark, as it conflates true consciousness with the simulation of consciousness.
That this is his popular legacy isn’t quite fair to Turing, who couldn’t have foreseen the ways that his ideas would be lifted from their original context. Turing knew that defining what it meant to “think” was tricky business. He had no intention to stake out a solution to the “hard problem of consciousness,” only to say that there were ways around traditional objections to intelligent machines by substituting one concept for another.
Neither was Turing the dispassionate logician many would assume. He elsewhere seems quite interested in the emotional and aesthetic dimensions of intelligence, imagining a different test where a contestant must 1) write a sonnet and 2) describe the reasoning behind it in order to demonstrate that they hold intelligence. He took sonnets as the highest form of literary excellence and humanity, a standard many humans probably couldn’t meet (including myself – my professor told me once that I “just didn’t quite get the form,” and he was right) but that, ironically, an LLM probably could in the eyes of most readers, at least in a blind test.
It’s also not fair to Lovelace that her words are mostly remembered through Turing’s quotation of her work.
Far from being a skeptic, Lovelace foresaw that the computational potential enabled by an “analytical engine” would someday bring new, unexpected outcomes, though she was wary of the “exaggerated ideas” that so often accompany technological speculation. Megan Ward’s chapter, “Victorian Fictions of Computational Creativity,” engages with Lovelace’s paper to a different end, contextualizing it in Victorian discourse about originality and creativity. I borrow liberally from her analysis here.
Victorian times were characterized by mechanical metaphors – what better way to describe the Industrial Revolution’s anxiety and exhilaration? Center to discussions of the industrial age was the word “origination,” which, in a time in which traditional crafts were being overtaken by automation, warranted a redefinition. No matter the fidelity with which they produced an object, these machines were incapable of “origination.”
Machines per se could not originate anything, but Lovelace echoed many of her time who argued that, when the capabilities of a machine were combined with human intelligence, a sort of symbiotic creativity was possible. Machines could handle complexity and combinations, while a human could decide what to pursue: by “distributing and combining the truths and the formulae of analysis… the relations and the nature of many subjects in that science are necessarily thrown into new lights, and more profoundly investigated.” The result is a form of creativity or originality that, as Ward puts it, escapes narrow, relativistically-defined attributes such as idiosyncrasy or surprise, expanding creativity into a collaborative process. If Turing’s definition of creativity is a mirror, Lovelace’s is a network.
For better or worse though, a half century of media telling implicit and explicit stories of Turing tests (Do Androids Dream of Electric Sheep, Her, Ex Machina – need I list the rest of the usual suspects?) have primed us to imagine that this is what intelligence is; that we are what intelligence looks like.
I see two dangers here: the first is what Erik Brynjolfsson terms the “Turing Trap,” where excessive focus on human-like artificial intelligence is actively harmful, as it invites replacement. Labor economists would probably agree – even in a better-case scenario where LLMs only have limited workplace relevance, “AI interns” might not threaten established professionals, but they make it a precarious time to be an apprentice, and many are worried that there will be a shortage of future trained workers.
The second is Goodhart’s Law, which is that “when a measure becomes a target, it ceases to be a good measure.” In this case, I believe that the (arguably successful) pursuit of Turing-tested AI limits our imaginations. Maybe if Turing had immortalized Lovelace’s words in a different way we would have optimized for something different, rather than so successfully simulating fluency. Language is only a dimension of intelligence, albeit a powerful one. I believe we privilege it at the expense of other types of intelligence, such as the distributed intelligence of octopods, where a bit of their mind is located in each tentacle; the network intelligence of slime molds that mirrors urban planning efforts; the systematic intelligence of vascular systems in cells and supply chains.
Or maybe thinking more creatively about creativity would change nothing; perhaps we would have landed here regardless. Maybe we’d find that Lovelace’s collaborative creativity is every bit the trap as Turing’s. For instance, it’s hard not to draw a line between Lovelace’s proposition and the calls for a “human-in-the-loop” that I keep seeing everywhere. “Reinforcement learning from human feedback,” the primary way that we train LLMs in the present, means that, given enough time, there’s a very fine line between augmenting a worker and training their replacement. It’s a problem I don’t see a way out of right now.
There’s an additional, eerie case for machine “origination” that I find compelling, if you’re willing to have an open mind about what creativity or intelligence is: Generative Adversarial Networks (GANs) are a class of machine learning, now mostly of interest to academics and artists, that were once the dominant approach to image generation. A GAN comprises two parts: a generator and a discriminator. The generator starts from noise and tries to produce a recognizable image, and the discriminator, trained on real images, tries to identify it as fake. As the two go back and forth, the generator becomes better at creating passable images (fooling the artificial judge — is this a sort of reverse Turing test?).
Where it gets interesting is when the two are untethered and the generator is allowed to create images without the discriminator to align it with human comprehension. The result is a representation of machine “dreaming,” where what we see is a purer manifestation of the latent space created during the generator’s training, or the opaque interiority of how a computer “thinks” on its own terms (if you want to get anthropomorphic about it). As with all things creative, it’s up to the viewer to determine what value is to be had here.

Jake Elwes’ “Latent Space”
Are GANs useful? Not particularly — they’ve been largely superseded by diffusion models like the ones behind DALL-E and Midjourney. Do they make interesting art? At least one artist thought so. And of course, no matter how mysterious or alien we currently find latent spaces, they’re still shaped by the deliberate decisions that train neural networks, and even purely “machine” outputs can defensibly be traced to human inputs. But thinking about them is a good exercise in imagination; a useful place to begin if you’re trying to determine what a machine can do uniquely well and what is seeded by a human collaborator.
I wish I could ask Lovelace what she makes of all this.

Meg Shriber (MBA ‘27) graduated from the University of California, Berkeley in 2022 with a degree in literature, and the University of Cambridge in 2024, where she wrote her MPhil dissertation on AI and creativity. Her first article, “Death of an Author, Birth of a Medium: Collaboration, Control, and Creativity in Machine-Generated Text” is forthcoming in Poetics Today. Meg is also a painter.




Comments