Skip to content

Essay

Engaged, but not drawn

Atención retenida, no atraída

A three-step claim about pictures and attention. Two steps rest on measurements nobody argues about. The one in the middle has never been tested.

· 9 min read

Read in

I have been carrying a three-step claim around, and in September 2026 I went looking for the evidence under it. It goes like this. An image reaches a reader before the words do. If that image has to be interpreted, it produces a stop. And the stop buys time to read what is next to it. One step turns out to rest on measurements I did not make. One returns a result I was not expecting. And the one in the middle is a question the field already has the instruments to answer, and has never pointed them at it.

Step one: the image gets in without asking.

The defensible version is not about speed. It is about access, and the two sides of it are measured. Adam Larson and Lester Loschky found that the gist of a scene can be categorized from a single fixation, and that peripheral vision carried more of that work than central vision did. Reading runs the other way. It requires the eye to land on each stretch of words in turn, and Marc Brysbaert's meta-analysis of 190 studies covering 18,573 participants puts silent reading of English prose at around 238 words a minute. Reading the asymmetry as access rather than as a race is my own interpretation. The measurements under it are not, and neither of them is mine. One of those is a rate you have to buy. The other is not. If the image arrives for free, the question is whether a free arrival can be turned into paid attention.

Step two: from a free arrival to a stop.

This is measurable, and it has never been measured under the conditions that would make it mean anything: a reader who did not agree to look, at an advertisement, with nobody telling them there was a figure to find.

What is measured, and measured well, is memory. Edward McQuarrie and David Glen Mick, in the Journal of Consumer Research in 2003, embedded advertisements in a 32-page magazine worth reading on its own, then varied whether people were told to process them or simply left to come across them. Advertisements built on a rhetorical figure were recalled more often and liked better, and the split matters more than the headline: visual figures worked whether or not anyone had been directed to look, while verbal figures only worked for readers who had been told to read. A clever line needs someone who agreed to read it. A picture that asks to be resolved does not.

That is not a stop. Everyone in that study was already looking, as they were in the work by Jaana Simola and colleagues showing that material requiring elaboration holds people longer once they are there. These are two different quantities: dwell time indexes retention, the timing of the first fixation indexes capture. And you cannot resolve a picture you have not yet looked at, so the interpretation that produces the memory is necessarily work done by someone who had already stopped.

Capture has been measured at scale on real advertising. Rik Pieters and Michel Wedel modeled 1,363 print advertisements for the Journal of Marketing in 2004, with eye-tracking data from more than 3,600 consumers, seen at the reader's own pace inside the magazines they belonged to. Nothing in that model turns on whether anything had to be worked out. So the field owns one of the two capture measures, whether an element gets fixated at all, and has never pointed it at this question. The nearest attempt, a study of visual complexity in 249 advertisements, separates complexity from how hard an advertisement is to understand. This question is about the second.

The nearest thing to an answer sits in a different literature. Geoffrey Loftus and Norman Mackworth reported in 1978 that twelve observers fixated improbable objects in a scene, an octopus on a farm, earlier, more often and for longer than probable ones. If that held, the step would be settled by accident. It does not hold cleanly. Several groups since have found the longer fixations and not the earlier ones, and Melissa Võ and John Henderson put the objection in their title: object-scene inconsistencies do not capture gaze. Using a paradigm built to test exactly that, they found no pull toward the inconsistent object from the periphery at all, and the eyes lingering on it only once they had already landed. The summary the field has settled on is that attention is engaged by incongruent objects rather than drawn to them. And an octopus on a farm is anomalous rather than resolvable. Nothing there asks a viewer to work something out and arrive somewhere.

Fig. 01

What each study's design let it record

Fig. 01. What each study's design let it record. No unbroken line crosses. Design only. No finding is shown here.
What they showedHow they watched
In the picturePicture matchedReal adNot told to lookWhether fixatedWhen first fixated
McQuarrie & Mick, 2003recordednot established in this piece's sourcesrecordedrecordednot metnot met
Simola et al., 2020recordednot established in this piece's sourcesrecordednot metnot metnot met
Pieters & Wedel, 2004not metnot metrecordedrecordedrecordednot met
Pieters et al., 2010not metnot metrecordednot established in this piece's sourcesnot established in this piece's sourcesnot established in this piece's sources
Loftus & Mackworth, 1978not metnot metnot metnot established in this piece's sourcesrecordedrecorded
Võ & Henderson, 2011not metnot metnot metnot established in this piece's sourcesrecordedrecorded
Konovalova & Petrova, 2023not metnot metrecordednot established in this piece's sourcesnot established in this piece's sourcesnot established in this piece's sources
The study that would settle itrequired by the test; not yet taken.required by the test; not yet taken.required by the test; not yet taken.required by the test; not yet taken.required by the test; not yet taken.required by the test; not yet taken.
  • recorded
  • not established in this piece's sources
  • not met
  • required by the test; not yet taken.

No unbroken line crosses. Design only. No finding is shown here.

I cannot prove it either, so I am putting it forward as a thesis.

A picture that can be resolved, and that offers a resolution worth reaching, holds the eye before the viewer has decided to give it anything. That is a claim I am proposing. It is not a result I am reporting.

Call it the stop. The name is neither mine nor new: stopping power has been in the Journal of Marketing since 2010, and the industry has been saying thumb-stopping, scroll-stopping and pattern interrupt for years. What I want to add is one clause. A stop is an unplanned halt in a self-directed sequence of looking, produced by what the content turns out to be rather than by its size, contrast or movement, and lasting long enough for the viewer to attempt a resolution. The middle clause is the one that matters, because none of the existing terms draws that line. Nobody has measured a stop so defined.

The idea they all circle is older than any of them. Viktor Shklovsky argued in 1917 that art makes its forms difficult on purpose, to lengthen a perception that habit has worn smooth. I have not read the original and take that from accounts of it. A century later, I could not find anyone who has measured it under the conditions this piece is about, which is not the same as saying nobody has.

Not for lack of a number: the industry already measures something it calls stopping. Thumb-stop rate is three-second video views over impressions, with hold rate behind it for what happens next. It is in production across a great many accounts, and in 2025 a study in the Journal of Advertising validated the family of measures it belongs to against mobile eye-tracking. But in a feed the video plays by itself, so what that number really counts is how long the ad stayed on the screen. It cannot tell a puzzle from a loud color, and that distinction is the whole of what I am adding.

It is also the distinction my own book does not name: that book describes how attention arrives parceled into glances, not what interrupts one.

Both ends of the chain are measured and only the middle is not. The thesis is the bridge between two measured facts, which is a better position than a hunch and a worse one than a finding.

What runs against it carries the same weight. The incongruity literature leans toward engaged rather than drawn. And effort by itself does nothing. Heping Xie, Zongkui Zhou and Qingqi Liu pooled twenty-five articles covering 3,135 participants for Educational Psychology Review in 2018, and found that degrading how a text looks had no effect on recall, at d = −0.01, while readers judged they had learned less and took longer. Difficulty is not the ingredient. Something to resolve is.

Step three: what the stop buys.

Anastasiia Konovalova and Tatiana Petrova, in the Journal of Eye Movement Research in 2023, tracked 53 Russian-speaking participants viewing 25 Russian advertising posters, eleven of which carried a pun in the headline. The general pattern of viewing did not change, and posters without puns were looked at significantly longer overall.

That reads like a flat no, and it is not. Dwell time rose specifically on the main text area for the posters carrying a pun. The authors read that as the pun being resolved, and they point out why it lands there: the second meaning of their puns was planted in the body copy, so the body copy is where the work had to be done. The total time result runs the other way, and the authors distrust it themselves, suggesting readers may have lingered on the pun-free posters hunting for puns that were not there.

What is left is narrower than a yes or a no: a figure can send attention to a specific place, if that is where you put the thing it resolves into.

The 2004 study puts a number on that from the other side. Pieters and Wedel also measured transfer, meaning how much attention each element passes to the others, and the element that transfers best is not the picture. It is the brand, measured, again, on pictures nobody had to work out.

Two limits keep this from settling the chain. The figure sat in the headline, which makes it the verbal case, and that is exactly where the 2003 study found the effect needs a willing reader. And the pictures were not controlled at all: the authors say the non-verbal side of their posters is very diverse, with content distributed unevenly across the two conditions. It went to the point that the authors had to enter the human figure as a covariate. Nobody varied whether the picture itself had to be worked out.

The test

Here is the experiment. Take matched pairs of advertisements where the only change is whether the pictorial requires a resolving step, a visual metaphor against a literal depiction of the same product, holding size, color, brand and copy constant. Then control the picture itself, which is where the existing work came apart: counterbalance so that each image serves in both roles across a normed set, or match the pairs on a computational salience model and report the match. Pretest that the figurative version really does demand a resolving step, and record whether each viewer resolved it, because without both a null result cannot be read. Show them under incidental, self-paced exposure inside a browsing task, never instructing anyone to look. Record time to first fixation and the probability of fixating at all, alongside dwell on the body copy. The result that decides the thesis is time to first fixation.

I am left with a thesis: that a picture which has to be worked out produces a stop, with a definition attached and an experiment waiting for somebody to run it. It was the step I was most confident about when I started. The nearest literature says attention is engaged by the unexpected rather than drawn to it, which is the sentence the thesis has to get past. It stays in the column marked proposed until that study is run, and everything needed to run it already exists.

Five questions before you approve a picture that has to be worked out

Five questions worth putting in a brief, or asking in the room before anyone signs off on an image with a puzzle in it.

  1. What exactly is the viewer meant to arrive at, and has anyone outside this room arrived there without help? In McQuarrie and Mick's 1999 study the effect on visual tropes weakened for people who did not share the cultural reference, and it was the pleasure of resolving them that went, not the comprehension.

  2. Is the figure in the picture or in the line? Visual figures held up with readers nobody had asked to look. Verbal ones only worked for readers who had already agreed to read.

  3. Are we buying difficulty or something to resolve? Degrading how a text looks returned no effect on recall across twenty-five studies and 3,135 participants, while costing readers time and confidence.

  4. If the figure resolves into something, where have we put that something? The one study that found attention moving toward the body copy found it moving to the place where the answer had been planted.

  5. And if you have an ad account: would you run it? Two creatives matched on everything except whether the picture resolves into something. Your own numbers will tell you which one kept more people on screen, though the platform decides who sees each one, so the two are not quite a fair race. What they cannot tell you at all is whether it was the puzzle or the louder picture, and that is the measurement nobody has taken.

What I could not verify

Shklovsky's 1917 essay, which I report from accounts of it rather than from the original; the full text of the 2025 validation study, which I could not obtain and describe only from its abstract; the full text of the 2004 attention study, which is closed and which I could not obtain, so that what I say about its design and its measures rests on peer-reviewed secondary accounts rather than on the paper itself; and whether the recall figure in the 2018 meta-analysis survives the check its critics ran on transfer, because nobody has run it.

Sources

  • Bruns, D., Kopka, J. F., Borgmann, L., Prior, S. & Langner, T. (2025). “Measuring Gaining and Holding Attention to Social Media Ads with Viewport Logging: A Validation Study Using Mobile Eye-Tracking.” Journal of Advertising, 54(5), 655–672. 10.1080/00913367.2025.2524186 Abstract only.

  • Brysbaert, M. (2019). “How many words do we read per minute? A review and meta-analysis of reading rate.” Journal of Memory and Language, 109, 104047. 10.1016/j.jml.2019.104047 190 studies, 18,573 participants.

  • Konovalova, A. & Petrova, T. (2023). “Pun processing in advertising posters: evidence from eye tracking.” Journal of Eye Movement Research, 16(3), 5. 10.16910/jemr.16.3.5 Open access. N = 53, two excluded; 25 Russian posters, 11 with puns; Saint Petersburg State University. PMID 38370527.

  • Larson, A. M. & Loschky, L. C. (2009). “The contributions of central versus peripheral vision to scene gist recognition.” Journal of Vision, 9(10), 6. 10.1167/9.10.6 Verified in full text.

  • Loftus, G. R. & Mackworth, N. H. (1978). “Cognitive determinants of fixation location during picture viewing.” Journal of Experimental Psychology: Human Perception and Performance, 4(4), 565–572. 10.1037/0096-1523.4.4.565 N = 12. The earlier-fixation result has a split replication record; see Henderson, Weeks & Hollingworth (1999), Journal of Experimental Psychology: Human Perception and Performance, 25, 210–228, and De Graef, Christiaens & d’Ydewalle (1990).

  • McQuarrie, E. F. & Mick, D. G. (1999). “Visual Rhetoric in Advertising: Text-Interpretive, Experimental, and Reader-Response Analyses.” Journal of Consumer Research, 26(1), 37–54. 10.1086/209549

  • McQuarrie, E. F. & Mick, D. G. (2003). “Visual and Verbal Rhetorical Figures under Directed Processing versus Incidental Exposure to Advertising.” Journal of Consumer Research, 29(4), 579–587. 10.1086/346252

  • Pieters, R. & Wedel, M. (2004). “Attention Capture and Transfer in Advertising: Brand, Pictorial, and Text-Size Effects.” Journal of Marketing, 68(2), 36–50. 10.1509/jmkg.68.2.36.27794 Infrared eye tracking, 1,363 print advertisements, 3,600+ consumers.

  • Pieters, R., Wedel, M. & Batra, R. (2010). “The Stopping Power of Advertising: Measures and Effects of Visual Complexity.” Journal of Marketing, 74(5), 48–60. 10.1509/jmkg.74.5.048 249 advertisements, eye tracking.

  • Shklovsky, V. (1917). “Art as Technique.” Not read in the original; reported from secondary accounts. Modern translations often render the title as “Art as Device”; the essay is dated 1917 and appeared in an anthology in 1919.

  • Simola, J., Kuisma, J. & Kaakinen, J. K. (2020). “Attention, memory and preference for direct and indirect print advertisements.” Journal of Business Research, 111, 249–261. 10.1016/j.jbusres.2019.06.028

  • Võ, M. L.-H. & Henderson, J. M. (2011). “Object-scene inconsistencies do not capture gaze: evidence from the flash-preview moving-window paradigm.” Attention, Perception, & Psychophysics, 73(6), 1742–1753. 10.3758/s13414-011-0150-6

  • Weissgerber, S. C., Brunmair, M. & Rummer, R. (2021). “Null and Void? Errors in Meta-analysis on Perceptual Disfluency and Recommendations to Improve Meta-analytical Reproducibility.” Educational Psychology Review, 33(3), 1221–1247. 10.1007/s10648-020-09579-1 Commentary.

  • Xie, H., Zhou, Z. & Liu, Q. (2018). “Null Effects of Perceptual Disfluency on Learning Outcomes in a Text-Based Educational Context: A Meta-Analysis.” Educational Psychology Review, 30(3), 745–771. 10.1007/s10648-018-9442-x

Llevo tiempo dándole vueltas a una afirmación de tres pasos, y en septiembre de 2026 fui a buscar la evidencia que la sostiene. Dice así: una imagen llega al lector antes que las palabras; si esa imagen hay que interpretarla, produce una parada; y la parada compra tiempo para leer lo que tiene al lado. De los tres pasos, uno se apoya en mediciones que no son mías, otro devuelve un resultado que no esperaba, y el del medio es una pregunta que el campo ya tiene instrumentos para responder, aunque nunca se los haya apuntado.

Paso uno: la imagen entra sin pedir permiso.

La versión defendible no habla de velocidad, sino de acceso, y los dos lados del acceso están medidos. Adam Larson y Lester Loschky encontraron que el sentido general de una escena se puede categorizar a partir de una sola fijación, y que la visión periférica cargaba con más de ese trabajo que la central. La lectura funciona al revés, porque exige que el ojo aterrice en cada tramo de palabras por turno: el metaanálisis de Marc Brysbaert, sobre 190 estudios con 18.573 participantes, sitúa la lectura silenciosa de prosa en inglés en unas 238 palabras por minuto. Leer la asimetría como acceso y no como carrera es mi propia interpretación; las mediciones que hay debajo no lo son, y ninguna de las dos es mía. Una de ellas es una velocidad que hay que comprar y la otra no lo es. Si la imagen llega gratis, la pregunta es si esa llegada gratuita se puede convertir en atención pagada.

Paso dos: de una llegada gratuita a una parada.

Esto se puede medir, y nunca se ha medido en las condiciones que le darían sentido: un lector que no aceptó mirar, frente a un anuncio, sin que nadie le haya avisado de que hay una figura que encontrar.

Lo que sí está medido, y bien medido, es la memoria. Edward McQuarrie y David Glen Mick, en el Journal of Consumer Research en 2003, insertaron anuncios en una revista de 32 páginas que valía la pena leer por sí sola, y luego variaron si a los participantes se les pedía procesarlos o si simplemente se los dejaba topárselos. Los anuncios construidos sobre una figura retórica se recordaron más y gustaron más, pero lo que importa no es el titular, sino la división: las figuras visuales funcionaron tanto con quienes habían recibido la indicación de mirar como con quienes no, mientras que las verbales solo funcionaron con los lectores a los que se les había pedido leer. Una línea ingeniosa necesita a alguien que haya aceptado leerla; una imagen que pide que la resuelvan, no.

Eso no es una parada. Todos los de ese estudio ya estaban mirando, igual que en el trabajo de Jaana Simola y sus colegas, que muestra que el material que exige elaboración retiene a la gente más tiempo una vez que ya está ahí. Son dos magnitudes distintas: el tiempo de permanencia es un índice de la retención, mientras que el momento de la primera fijación lo es de la captura. Y como no se puede resolver una imagen que todavía no se ha mirado, la interpretación que produce ese recuerdo es necesariamente trabajo de alguien que ya se había parado.

La captura sí se ha medido a escala y sobre publicidad real. Rik Pieters y Michel Wedel modelaron 1.363 anuncios impresos para el Journal of Marketing en 2004, con datos de seguimiento ocular de más de 3.600 consumidores que los vieron a su propio ritmo y dentro de las revistas a las que pertenecían. Nada en ese modelo depende de que hubiera algo que resolver. Así que el campo tiene una de las dos medidas de captura, la de si un elemento llega a recibir una fijación, y nunca la ha apuntado a esta pregunta. El intento más cercano, un estudio de complejidad visual sobre 249 anuncios, separa la complejidad de lo difícil que resulta entender un anuncio, y es de lo segundo de lo que va esta pregunta.

Lo más parecido a una respuesta está en otra literatura. Geoffrey Loftus y Norman Mackworth informaron en 1978 de que doce observadores fijaban los objetos improbables de una escena, un pulpo en una granja, antes, más veces y durante más tiempo que los probables. Si eso se sostuviera, el paso quedaría resuelto por accidente. No se sostiene limpiamente. Varios grupos han encontrado desde entonces las fijaciones más largas pero no las más tempranas, y Melissa Võ y John Henderson pusieron la objeción en su título: las inconsistencias entre objeto y escena no capturan la mirada. Con un paradigma construido para probar exactamente eso, no encontraron ningún tirón hacia el objeto inconsistente desde la periferia, y los ojos solo se demoraban en él una vez que ya habían aterrizado. El resumen en el que el campo se ha asentado es que la atención queda retenida por los objetos incongruentes en vez de ser atraída hacia ellos. Y un pulpo en una granja es anómalo, no algo que se resuelva: ahí no hay nada que le pida a quien mira trabajar algo y llegar a alguna parte.

Fig. 01

Lo que el diseño de cada estudio permitió registrar

Fig. 01. Lo que el diseño de cada estudio permitió registrar. Ninguna línea continua cruza. Solo diseño. Aquí no se muestra ningún hallazgo.
Qué mostraronCómo observaron
En la imagenImagen pareadaAnuncio realSin pedir que mirenSi hubo fijaciónCuándo la primera
McQuarrie & Mick, 2003registradano establecida en las fuentes de esta piezaregistradaregistradano cumplidano cumplida
Simola et al., 2020registradano establecida en las fuentes de esta piezaregistradano cumplidano cumplidano cumplida
Pieters & Wedel, 2004no cumplidano cumplidaregistradaregistradaregistradano cumplida
Pieters et al., 2010no cumplidano cumplidaregistradano establecida en las fuentes de esta piezano establecida en las fuentes de esta piezano establecida en las fuentes de esta pieza
Loftus & Mackworth, 1978no cumplidano cumplidano cumplidano establecida en las fuentes de esta piezaregistradaregistrada
Võ & Henderson, 2011no cumplidano cumplidano cumplidano establecida en las fuentes de esta piezaregistradaregistrada
Konovalova & Petrova, 2023no cumplidano cumplidaregistradano establecida en las fuentes de esta piezano establecida en las fuentes de esta piezano establecida en las fuentes de esta pieza
El estudio que lo zanjaríaexigida por la prueba; aún no tomada.exigida por la prueba; aún no tomada.exigida por la prueba; aún no tomada.exigida por la prueba; aún no tomada.exigida por la prueba; aún no tomada.exigida por la prueba; aún no tomada.
  • registrada
  • no establecida en las fuentes de esta pieza
  • no cumplida
  • exigida por la prueba; aún no tomada.

Ninguna línea continua cruza. Solo diseño. Aquí no se muestra ningún hallazgo.

Yo tampoco puedo demostrarlo, así que lo presento como una tesis.

Una imagen que se puede resolver, y que ofrece una resolución que vale la pena alcanzar, retiene el ojo antes de que quien mira haya decidido darle nada. Eso es una afirmación que propongo, no un resultado que comunico.

Llámalo la parada. El nombre no es mío ni es nuevo: stopping power está en el Journal of Marketing desde 2010, y la industria lleva años diciendo thumb-stopping, scroll-stopping y pattern interrupt. Lo que quiero añadir es una cláusula.

Una parada es una detención no planeada dentro de una secuencia de mirada que dirige uno mismo, producida por lo que el contenido resulta ser y no por su tamaño, su contraste o su movimiento, y que dura lo suficiente para que quien mira intente una resolución.

La cláusula del medio es la que importa, porque ninguno de los términos que ya existen traza esa línea. Nadie ha medido una parada definida así.

La idea que todos ellos rodean es más vieja que cualquiera de ellos. Viktor Shklovsky sostuvo en 1917 que el arte vuelve difíciles sus formas a propósito, para alargar una percepción que la costumbre ha desgastado. No he leído el original y lo tomo de fuentes secundarias. Un siglo después, no encontré a nadie que lo haya medido en las condiciones de las que va esta pieza, lo cual no es lo mismo que decir que nadie lo haya hecho.

Y no es por falta de una cifra: la industria ya mide algo que llama detenerse. El thumb-stop rate son las visualizaciones de video de tres segundos sobre las impresiones, con el hold rate detrás para lo que pasa después. Está en producción en muchísimas cuentas, y en 2025 un estudio del Journal of Advertising validó, contra seguimiento ocular en móvil, la familia de medidas a la que pertenece. Pero en un feed el video se reproduce solo, así que lo que esa cifra cuenta en realidad es cuánto tiempo el anuncio se quedó en pantalla. No puede distinguir un acertijo de un color chillón, y esa distinción es todo lo que estoy añadiendo.

Es también la distinción que mi propio libro no nombra: ese libro describe cómo la atención llega repartida en vistazos, no qué interrumpe un vistazo.

Los dos extremos de la cadena están medidos y solo el del medio no lo está. La tesis es el puente entre dos hechos medidos, que es una posición mejor que una intuición y peor que un hallazgo.

Lo que va en su contra tiene el mismo peso. La literatura de la incongruencia se inclina por retenida y no por atraída. Y el esfuerzo por sí solo no hace nada: Heping Xie, Zongkui Zhou y Qingqi Liu agruparon veinticinco artículos con 3.135 participantes para Educational Psychology Review en 2018, y encontraron que degradar el aspecto de un texto no tuvo ningún efecto sobre el recuerdo, con d = −0.01, mientras que los lectores juzgaron haber aprendido menos y tardaron más. La dificultad no es el ingrediente. Algo que resolver, sí.

Paso tres: qué compra la parada.

Anastasiia Konovalova y Tatiana Petrova, en el Journal of Eye Movement Research en 2023, siguieron a 53 participantes rusohablantes mientras veían 25 carteles publicitarios rusos, once de los cuales llevaban un juego de palabras en el titular. El patrón general de mirada no cambió, y los carteles sin juego de palabras se miraron significativamente más tiempo en total.

Eso se lee como un no rotundo y no lo es. El tiempo de permanencia subió específicamente sobre el área de texto principal en los carteles que llevaban juego de palabras. Las autoras lo leen como el juego de palabras resolviéndose, y señalan por qué aterriza ahí: el segundo significado de sus juegos de palabras estaba plantado en el cuerpo de texto, así que el cuerpo de texto es donde había que hacer el trabajo. El resultado del tiempo total va en la dirección contraria, y las propias autoras desconfían de él, porque apuntan que los lectores pueden haberse demorado en los carteles sin juego de palabras buscando juegos que no estaban.

Lo que queda es más estrecho que un sí o un no: una figura puede mandar la atención a un sitio concreto, siempre que sea ahí donde hayas puesto aquello en lo que se resuelve.

El estudio de 2004 le pone una cifra a eso desde el otro lado. Pieters y Wedel midieron también la transferencia, es decir, cuánta atención le pasa cada elemento a los demás, y el elemento que mejor transfiere no es la imagen. Es la marca, medida, otra vez, sobre imágenes que nadie tenía que resolver.

Dos límites impiden que esto zanje la cadena. La figura estaba en el titular, lo que la convierte en el caso verbal, y es exactamente ahí donde el estudio de 2003 encontró que el efecto necesita un lector dispuesto. Y las imágenes no estaban controladas en absoluto: las autoras dicen que el lado no verbal de sus carteles es muy diverso, con el contenido distribuido de forma desigual entre las dos condiciones, hasta el punto de que tuvieron que introducir la figura humana como covariable. Nadie varió si había que resolver la imagen misma.

La prueba

Aquí está el experimento. Toma pares equiparados de anuncios en los que lo único que cambie sea si lo pictórico exige un paso de resolución: una metáfora visual frente a una representación literal del mismo producto, manteniendo constantes el tamaño, el color, la marca y el texto. Después controla la imagen misma, que es donde el trabajo existente se deshizo, y hazlo de una de estas dos formas: contrabalancea de modo que cada imagen sirva en los dos papeles a lo largo de un conjunto normado, o equipara los pares con un modelo computacional de saliencia y declara la equiparación. Comprueba de antemano que la versión figurativa exige de verdad un paso de resolución, y registra si cada persona lo resolvió, porque sin las dos cosas un resultado nulo no se puede interpretar. Muestra los anuncios en exposición incidental y al ritmo del propio lector, dentro de una tarea de navegación, sin pedirle nunca a nadie que mire. Registra el tiempo hasta la primera fijación y la probabilidad de que el elemento reciba fijación alguna, junto con la permanencia sobre el cuerpo de texto. El resultado que decide la tesis es el tiempo hasta la primera fijación.

Me queda una tesis: que una imagen que hay que resolver produce una parada, con una definición puesta y un experimento esperando a que alguien lo lleve a cabo. Era el paso del que estaba más seguro cuando empecé. La literatura más cercana dice que la atención queda retenida por lo inesperado en vez de ser atraída hacia ello, y esa es la frase que la tesis tiene que superar. Se queda en la columna de lo propuesto hasta que ese estudio se haga, y todo lo que hace falta para hacerlo ya existe.

Cinco preguntas antes de aprobar una imagen que hay que resolver

Cinco preguntas que vale la pena poner en un brief, o hacer en la sala antes de que alguien firme una imagen con un acertijo dentro.

  1. ¿A qué tiene que llegar exactamente quien mira, y ha llegado ahí alguien de fuera de esta sala sin ayuda? En el estudio de 1999 de McQuarrie y Mick el efecto sobre los tropos visuales se debilitó en la gente que no compartía la referencia cultural, y lo que se perdió fue el placer de resolverlos, no la comprensión.

  2. ¿La figura está en la imagen o en la línea? Las figuras visuales aguantaron con lectores a los que nadie había pedido mirar, mientras que las verbales solo funcionaron con lectores que ya habían aceptado leer.

  3. ¿Estamos comprando dificultad o algo que resolver? Degradar el aspecto de un texto no devolvió ningún efecto sobre el recuerdo a lo largo de veinticinco estudios y 3.135 participantes, y además les costó a los lectores tiempo y confianza.

  4. Si la figura se resuelve en algo, ¿dónde hemos puesto ese algo? El único estudio que encontró la atención moviéndose hacia el cuerpo de texto la encontró moviéndose al sitio donde estaba plantada la respuesta.

  5. Y si tienes una cuenta de anuncios, ¿lo lanzarías? Dos creatividades equiparadas en todo menos en si la imagen se resuelve en algo. Tus propias cifras te dirán cuál retuvo a más gente en pantalla, aunque la plataforma decide quién ve cada una, así que las dos no compiten en una carrera del todo justa. Lo que no te pueden decir en absoluto es si fue el acertijo o la imagen más chillona, y esa es la medición que nadie ha tomado.

Lo que no pude verificar

El ensayo de Shklovsky de 1917, del que doy cuenta a partir de fuentes secundarias y no del original; el texto completo del estudio de validación de 2025, que no pude conseguir y que describo solo a partir de su resumen; el texto completo del estudio de atención de 2004, que es cerrado y que tampoco pude conseguir, de modo que lo que digo de su diseño y de sus medidas se apoya en fuentes secundarias revisadas por pares y no en el artículo mismo; y si la cifra de recuerdo del metaanálisis de 2018 sobrevive a la comprobación que sus críticos hicieron sobre la transferencia, porque nadie la ha hecho.

Fuentes

  • Bruns, D., Kopka, J. F., Borgmann, L., Prior, S. & Langner, T. (2025). “Measuring Gaining and Holding Attention to Social Media Ads with Viewport Logging: A Validation Study Using Mobile Eye-Tracking.” Journal of Advertising, 54(5), 655–672. 10.1080/00913367.2025.2524186 Abstract only.

  • Brysbaert, M. (2019). “How many words do we read per minute? A review and meta-analysis of reading rate.” Journal of Memory and Language, 109, 104047. 10.1016/j.jml.2019.104047 190 studies, 18,573 participants.

  • Konovalova, A. & Petrova, T. (2023). “Pun processing in advertising posters: evidence from eye tracking.” Journal of Eye Movement Research, 16(3), 5. 10.16910/jemr.16.3.5 Open access. N = 53, two excluded; 25 Russian posters, 11 with puns; Saint Petersburg State University. PMID 38370527.

  • Larson, A. M. & Loschky, L. C. (2009). “The contributions of central versus peripheral vision to scene gist recognition.” Journal of Vision, 9(10), 6. 10.1167/9.10.6 Verified in full text.

  • Loftus, G. R. & Mackworth, N. H. (1978). “Cognitive determinants of fixation location during picture viewing.” Journal of Experimental Psychology: Human Perception and Performance, 4(4), 565–572. 10.1037/0096-1523.4.4.565 N = 12. The earlier-fixation result has a split replication record; see Henderson, Weeks & Hollingworth (1999), Journal of Experimental Psychology: Human Perception and Performance, 25, 210–228, and De Graef, Christiaens & d’Ydewalle (1990).

  • McQuarrie, E. F. & Mick, D. G. (1999). “Visual Rhetoric in Advertising: Text-Interpretive, Experimental, and Reader-Response Analyses.” Journal of Consumer Research, 26(1), 37–54. 10.1086/209549

  • McQuarrie, E. F. & Mick, D. G. (2003). “Visual and Verbal Rhetorical Figures under Directed Processing versus Incidental Exposure to Advertising.” Journal of Consumer Research, 29(4), 579–587. 10.1086/346252

  • Pieters, R. & Wedel, M. (2004). “Attention Capture and Transfer in Advertising: Brand, Pictorial, and Text-Size Effects.” Journal of Marketing, 68(2), 36–50. 10.1509/jmkg.68.2.36.27794 Infrared eye tracking, 1,363 print advertisements, 3,600+ consumers.

  • Pieters, R., Wedel, M. & Batra, R. (2010). “The Stopping Power of Advertising: Measures and Effects of Visual Complexity.” Journal of Marketing, 74(5), 48–60. 10.1509/jmkg.74.5.048 249 advertisements, eye tracking.

  • Shklovsky, V. (1917). “Art as Technique.” Not read in the original; reported from secondary accounts. Modern translations often render the title as “Art as Device”; the essay is dated 1917 and appeared in an anthology in 1919.

  • Simola, J., Kuisma, J. & Kaakinen, J. K. (2020). “Attention, memory and preference for direct and indirect print advertisements.” Journal of Business Research, 111, 249–261. 10.1016/j.jbusres.2019.06.028

  • Võ, M. L.-H. & Henderson, J. M. (2011). “Object-scene inconsistencies do not capture gaze: evidence from the flash-preview moving-window paradigm.” Attention, Perception, & Psychophysics, 73(6), 1742–1753. 10.3758/s13414-011-0150-6

  • Weissgerber, S. C., Brunmair, M. & Rummer, R. (2021). “Null and Void? Errors in Meta-analysis on Perceptual Disfluency and Recommendations to Improve Meta-analytical Reproducibility.” Educational Psychology Review, 33(3), 1221–1247. 10.1007/s10648-020-09579-1 Commentary.

  • Xie, H., Zhou, Z. & Liu, Q. (2018). “Null Effects of Perceptual Disfluency on Learning Outcomes in a Text-Based Educational Context: A Meta-Analysis.” Educational Psychology Review, 30(3), 745–771. 10.1007/s10648-018-9442-x