AI-generated food images are increasingly appearing in restaurant, café and brand marketing, but their warped sandwiches, tangled noodles and hole-ridden burritos are often more likely to repel customers than whet their appetites. Experts say the disturbing results stem from a combination of technical limitations, flawed training material and the way people instinctively recognise potential signs of spoiled or contaminated food.
Many image generators use diffusion models, which begin with visual noise and gradually refine it into a finished picture. Chris Russell, a professor of AI, government and policy at the University of Oxford, said these systems recover broad shapes before adding fine details.
That order of operations can produce an image whose basic structure is already wrong when the texture is applied. The result may be a burger with impossible layers, prawns fused into dough or a pastry covered in details that belong somewhere else.
“This means that initially coarse structures are recovered first with fine texture details” coming at the end, Russell said.
The problem is especially visible in foods made up of narrow, repeating forms. Giovanbattista Califano, a behavioural scientist at the University of Naples Federico II, said diffusion models struggle with strands, noodles and tendrils because they cannot reliably determine where such shapes should end or what they should connect to.
“Diffusion models are notoriously weak at generating thin, continuous, terminating structures,” Califano said. Strands may therefore spread across the image, while bubbles, seeds and other repeated patterns can escape the boundaries of the food they are meant to represent.
Roland Meyer, a professor of digital cultures and arts at the University of Zurich, said image generators reproduce the appearance of objects without understanding what those objects are or how they exist in the physical world. A model can learn visual associations linked to a sandwich or burrito, but it does not know their ingredients, construction or purpose.
That lack of context also helps explain why some synthetic food resembles building materials. Michael Cook, a senior lecturer in computer science at King’s College London, said a cracked-concrete texture might look ordinary in an architectural image but becomes unsettling when presented as ice cream or a burger.
Training data can intensify the effect. Food photography commonly uses glossy lighting, strong contrast, saturated colours and exaggerated shapes, while unusual images often attract more attention online than pictures of everyday meals. Simon Colton, a professor of computational creativity, games and artificial intelligence at Queen Mary University of London, said a model may encounter more striking or bizarre examples than ordinary food.
Cook also said AI-generated material is increasingly being used to train other systems. Repeatedly learning from synthetic images can contribute to “model collapse”, in which outputs become visually less varied and develop increasingly similar distortions.
Human reactions then complete the process. The uncanny valley is not limited to faces: people are highly sensitive to food that appears capable of carrying parasites, toxins or disease. Califano said the effect can be particularly visceral because disgust is thought to have evolved partly as a protective response.
Worm-like strands, clusters of holes, unnatural colours and contaminated-looking textures can trigger those warnings simultaneously. The images may contain recognisable features of food, but the brain detects that the overall object could not exist safely in the real world.
Vague prompts and poor-quality source images can make matters worse, particularly when a small generated picture is enlarged for advertising. Imperfections that might go unnoticed at low resolution become prominent, leaving brands with promotional material that looks less like a carefully prepared meal and more like an accidental digital mutation.
