The Senses Were Always the First Interface
Most experiential technology doesn't fail because the graphics aren't good enough. It fails because it lies to a nervous system that has spent several hundred million years learning to catch liars. What neuroscience — sensor fusion, latency thresholds, predictive processing — teaches a product roadmap, and why coherence, not fidelity, is the real product.

Nobody in a product review ever says “the latency is 140 milliseconds.” What they say is “something’s off.” They can’t point at it. They just know, in the specific and slightly nauseated way a body knows things before a mind has caught up, that the thing in front of them is lying.
Here’s the thesis, stated early because I’ve been told, rightly, that I bury these too deep: most experiential technology doesn’t fail because the graphics aren’t good enough. It fails because it lies to a nervous system that has spent several hundred million years learning to catch liars. Everything else in this piece is really just unpacking that sentence with better neuroscience than it usually gets.
Your brain has been doing sensor fusion since before it had a name for it
Here is the bit that gets skipped in most conversations about AR, IoT, robotics, and the whole broad church now filed under “experiential technology”: the brain was never a passive receiver of one clean channel of reality. It is, and has always been, a fusion engine — constantly reconciling vision, touch, proprioception (the sense of where your own limbs are in space), and the vestibular system (balance and motion, sitting quietly in your inner ear, mostly unthanked) into a single, coherent, felt experience. It does this so seamlessly that you’ve never once had to think about it, which is exactly why it’s worth thinking about now.
Two experiments make this uncomfortably vivid. The first is the McGurk effect: play someone audio of the syllable “ba” while showing them a video of a mouth mouthing “ga,” and most people will report hearing “da” — a sound that was never actually spoken. The eye edited the ear in real time, without permission, and without the listener noticing the edit happened at all. The second is the rubber hand illusion: stroke someone’s hidden real hand and a visible rubber hand on the table in perfect synchrony, and within a couple of minutes most people report a distinct sensation that the rubber hand is theirs — some will flinch if you threaten it with a hammer. The brain didn’t check the paperwork. It checked the timing.

That’s the whole game, right there. The brain doesn’t experience “real” and “synthetic” as separate categories to be adjudicated. It experiences coherence — signals that arrive correlated, on time, and mutually consistent — and it assembles whatever produces that coherence into a single, unquestioned experience of “this is happening to me.” It has never once cared whether the source was biological or synthetic. It only cares whether the story holds together.
This is either the most exciting or the most alarming fact in the entire experiential technology industry, depending on which side of the build you sit on. Exciting, because it means genuine presence — the felt sense that a digital layer is actually part of your physical world — is achievable with far less computational brute force than people assume, if you get the fusion right. Alarming, because the brain that assembled the rubber hand illusion in ninety seconds is exactly as good at un-assembling it the moment one channel contradicts another. Which brings us to the four things this actually teaches a product roadmap.
Four things neuroscience can teach a product roadmap
One: the nervous system doesn’t grade on realism, it grades on correlation. A cartoonish AR overlay that moves with zero perceptible lag will feel more “real” — more trusted, more present — than a photorealistic one that lags by 150 milliseconds. Fidelity is not the primary lever. Timing is. Most experiential-technology budgets are spent backwards: enormous sums on visual polish, and comparatively little on the plumbing that keeps every sensory channel synchronized. That ratio is almost always wrong.
Two: latency is the tell, and it has a number. Research on visually induced motion sickness in VR consistently points to roughly 20 milliseconds as the rough ceiling for motion-to-photon latency before a mismatch becomes detectable, and to something in the neighborhood of 100 milliseconds or beyond before it reliably produces the specific queasy wrongness engineers call cybersickness — your vestibular system reporting one motion while your eyes report another, and your brain, unable to reconcile the two, treating the conflict as a plausible symptom of poisoning. This is not a UX nicety. It is a hard physiological threshold, and it applies well beyond headsets: a smart-shelf sensor that updates the app three seconds after the physical stock actually moved, a wearable haptic cue that fires a beat after the on-screen action it’s meant to accompany — these are the same category of failure at a slower clock speed. The body notices the gap before the mind can name it.
Three: the illusion is real, but only while nothing contradicts it. The rubber hand illusion is robust — until the experimenter strokes the real and fake hands out of sync, at which point it collapses immediately and completely, and does not rebuild on the same trust the second time. Brand and product experience work identically. A phygital retail experience where the app says one thing and the shelf says another; an IoT device whose physical indicator light disagrees with its dashboard state; a robot whose movement doesn’t match the sound it’s making — each of these is one asynchronous stroke away from breaking an illusion of coherence that took real engineering to build and takes one bad sync to undo. The trust doesn’t degrade gracefully. It falls off a cliff, and the second attempt starts from further behind than the first one did.
Four: mismatch doesn’t create the fault, it reveals it. This is the lesson worth sitting with longest, because it explains why “bolted-on” digital layers so reliably feel worse than no digital layer at all. The rubber hand illusion doesn’t fail because synchrony is hard to engineer — it fails because the underlying setup was never actually integrated, just placed next to something real and hoped into coherence. Most failed AR features, most abandoned smart-home devices, most “digital transformation” retail pilots follow the identical pattern: a genuinely separate system, decorated to look connected, that a nervous system built for exactly this kind of detection work sniffs out within seconds. The mismatch isn’t the problem. It’s the diagnosis.
The cost has actually changed, which is the useful part
There’s an obvious objection here, and it’s the correct one: building genuinely synchronized, low-latency, multi-channel experiences used to require budgets that only the largest consumer tech companies could justify. Sensor fusion, spatial tracking, real-time device orchestration — this was, until quite recently, a hardware-lab problem, not a product-team problem.
That’s changed faster than most roadmaps have caught up with. Spatial computing SDKs, consumer-grade IoT platforms, and increasingly capable on-device inference have collapsed a huge amount of what used to be bespoke systems engineering into configuration. A team of a handful of people — the kind of team a scrappy CMO or a solo operator is actually working with — can now build a coherent, low-latency phygital experience that would have required a hardware division a decade ago. This is precisely the same shift as the memory-and-follow-through tooling I mentioned in the last piece: not a replacement for judgement, but an enormous drop in the price of the infrastructure that used to gate whether a small team could even attempt the ambitious version of an idea. The lattice — the synchronization, the fusion, the coherence-holding plumbing — got radically cheaper to build. That doesn’t make the neuroscience easier to satisfy. It makes it easier to afford satisfying it.
For marketers specifically, this is worth saying plainly: an experiential campaign is not a creative problem wearing a technical costume. It is a systems problem — get the timing, the cross-channel consistency, and the sensor-to-signal latency right, and the creative genuinely lands as presence rather than gimmick. Get it wrong, and no amount of creative brilliance survives contact with a nervous system that clocked the 200-millisecond gap before the first person in the room said a word out loud.
The process that actually builds the lattice
Knowing the neuroscience tells you what to check for. It doesn’t tell you how to build it, and this is where design thinking earns its keep rather than its reputation as a workshop word stuck on Post-its. The five-stage version — empathize, define, ideate, prototype, test — is, not coincidentally, structured almost exactly like the coherence problem itself. You cannot design synchronized cross-channel experience from a whiteboard; you have to empathize your way into the actual sensory and situational constraints of the person standing in the store, wearing the headset, or holding the phone at an odd angle in bad light, before you can define what “coherent” even means for that specific context. Ideation without that grounding produces the bolted-on layer from lesson four — a feature that looks connected in the deck and gets caught within seconds by the nervous system it was built for. Prototyping and testing exist specifically to catch the asynchronous stroke before a paying customer does. Skip a stage, and you get the rubber hand illusion built by someone who never watched what happens when the timing is off.
The other piece worth building into the roadmap, particularly for anyone making the investment case, is honesty about where a given technology actually sits on the Gartner Hype Cycle rather than where the pitch deck says it sits. Every wave of experiential technology — AR glasses, IoT-connected retail, embodied robotics — moves through the same five stations: the innovation trigger, the peak of inflated expectations, the trough of disillusionment, the slope of enlightenment, and eventually, for the technologies that make it, a plateau of productivity. Most of the “bolted-on” failures in lesson four aren’t engineering failures at all. They’re timing failures — a team building plateau-of-productivity execution quality for a technology that’s still somewhere in the trough, spending real budget on polish for infrastructure that isn’t stable enough yet to be polished. The useful discipline isn’t avoiding the trough. It’s knowing which station you’re actually standing in before you decide how much coherence-engineering the budget can justify.
Four things a good activation should actually do
Put the neuroscience and the process together and you get a fairly clean test for any proposed product activation or customer experience: does it educate, does it gather first-party data honestly, does it empower the person going through it, and does it accelerate their path to value. Most activations manage one of these. The good ones manage at least two without the other channels contradicting each other.

Educate, done well, looks like Sephora’s or IKEA’s AR try-before-you-buy layers — the digital overlay isn’t decoration, it’s teaching the customer something true about fit or scale that the physical product alone couldn’t convey fast enough. Done badly, it’s an explainer video bolted onto a product page that nobody watches because nothing about it responds to what the viewer is actually looking at.
Gather first-party data, in a post-cookie world, is arguably the most economically important of the four, and the most easily botched. Connected packaging — NFC tags, QR-linked warranty registration, app-paired hardware — can capture consented, first-party signal directly at the point of physical interaction, which is exactly the data third-party tracking can no longer reliably supply. Starbucks’ app-and-loyalty flywheel is the textbook version: every physical order becomes a first-party data point that improves the next digital recommendation, and vice versa, with the two channels never once disagreeing with each other. The failure mode is the loyalty scheme that demands the data up front and gives nothing coherent back — which is lesson three again, just wearing a marketing hat.
Empower shows up as self-serve configurators, real-time personalization, and community-driven platforms that hand the customer genuine agency rather than a curated illusion of it — Peloton’s connected hardware plus live and on-demand content is as much an empowerment mechanic as a fitness product, because the physical exertion and the digital leaderboard are reporting the same coherent story back to the user in real time. The failure mode is a “personalization” engine that recommends based on stale data — the digital channel confidently reporting something the physical relationship has already outgrown.
Accelerate is the one systems architects will recognize fastest, because it’s the one most directly gated by the latency numbers from lesson two: real-time inventory synced to an AR fitting or placement view, checkout that doesn’t ask the customer to re-enter what a physical tap-to-pay already told it. Every millisecond of gap between physical action and digital acknowledgment is friction the customer feels as doubt, whether or not they could name it. Acceleration, in this frame, isn’t a feature — it’s simply what coherence feels like from the outside.
None of these four is achievable by the marketing function alone, and that’s rather the point of this whole piece. Educate, gather, empower, accelerate are marketing outcomes sitting directly on top of an engineering substrate — the same timing, sync, and cross-channel consistency work from the neuroscience section, aimed outward at a customer instead of inward at a headset. Get the substrate wrong and no campaign brief survives contact with it.
The next interface isn’t visual. It’s predictive.
There’s a layer underneath everything covered so far, and it changes who this piece is actually about. The brain doesn’t just fuse the senses it’s currently receiving. Increasingly, the dominant account in neuroscience — predictive processing, sometimes called the brain as prediction engine — holds that perception isn’t built from the bottom up out of raw sensory data at all. The brain runs a running forecast of what it expects to sense next, and what you consciously experience is largely that forecast, corrected by whatever prediction error actually arrives. You don’t see the world and then understand it. You predict the world and then get quietly notified when it disagrees with you.

This is, in fact, the more complete reading of the two experiments from earlier. The McGurk effect isn’t just the eye overruling the ear — it’s the brain’s prediction of the most probable syllable, given both channels, simply winning the vote. The rubber hand illusion isn’t just proprioception getting hijacked — it’s the brain updating its predictive model of “where is my hand” because a rubber hand, stroked in perfect sync, became the better prediction. Coherence, all along, was really just prediction being confirmed rather than violated. Mismatch was prediction being violated, which is precisely why it recruits attention so violently and so fast — a prediction error is, biologically speaking, the single most urgent thing a nervous system can notice.
This is why a conversational AI that remembers what you told it three exchanges ago feels intelligent long before it’s doing anything especially clever. It isn’t the answer that creates the impression. It’s the continuity — the fact that the next response confirms your prediction that it understood the last one. The same principle, run in reverse, explains why a robot with flawless movement can still read as uncanny, why a technically correct digital assistant can still feel incompetent, and why a customer journey with every individual component working can still feel broken. The fault, in each case, usually isn’t in the capability. It’s in the prediction. The user expected one continuous story. The system delivered a correct but disconnected fragment of one, and the nervous system — the same one that caught the rubber hand’s one asynchronous stroke — noticed immediately.
Sensory mismatch reveals faults in physical experience. Expectation mismatch reveals faults in cognitive experience, by exactly the same mechanism, one level up the stack. Which means the next generation of interfaces — AI-driven ones especially — won’t ultimately be judged by pixels, sensors, or displays at all. They’ll be judged by how well they maintain a coherent model of what the human on the other end believes is currently happening. That’s a genuinely different design brief than the one most teams think they’re working from.
The point isn’t the illusion
The remarkable thing about experiential technology is not that it can fool the eye. Plenty of things fool the eye — a painting does that, and has for tens of thousands of years, without anyone mistaking it for a systems achievement. The remarkable thing is when it convinces the whole nervous system at once — vision, touch, balance, proprioception, and now expectation itself — all reporting the same coherent story, without ever getting caught contradicting itself.
That’s not a design win. It’s an engineering one. It is the direct, unglamorous result of treating latency, synchrony, and cross-channel consistency — and, increasingly, continuity of context — as the actual product, and treating the visible layer as the thin, final coat of paint on top of plumbing that either holds together or doesn’t.
The companies that win the next decade of experiential technology will not be the ones with the most impressive demos. They will be the ones that understand a simple truth: humans do not experience technology channel by channel. We experience it as a single story. The moment that story becomes incoherent, trust evaporates. The moment it holds together, technology disappears and experience remains.
That has been true for hundreds of millions of years. The senses were merely the first interface through which we learned it.
“Perception is not something that happens to us. It is something we do.” — Alva Noë, Action in Perception (2004)
Further reading
Core neuroscience
- McGurk, H., & MacDonald, J. (1976). Hearing lips and seeing voices. Nature, 264, 746–748.
- Botvinick, M., & Cohen, J. (1998). Rubber hands ‘feel’ touch that eyes see. Nature, 391, 756.
- Stein, B. E., & Stanford, T. R. (2008). Multisensory integration: current issues from the perspective of the single neuron. Nature Reviews Neuroscience.
Predictive processing
- Clark, A. (2013). Whatever next? Predictive brains, situated agents, and the future of cognitive science. Behavioral and Brain Sciences.
- Friston, K. (2010). The free-energy principle: a unified brain theory? Nature Reviews Neuroscience.
Presence and VR
- Slater, M. (2009). Place illusion and plausibility can lead to realistic behaviour in immersive virtual environments. Philosophical Transactions of the Royal Society B.
- LaViola, J. J. (2000). A discussion of cybersickness in virtual environments. ACM SIGCHI Bulletin.
Human factors
- Parasuraman, R., Sheridan, T. B., & Wickens, C. D. (2000). A model for types and levels of human interaction with automation. IEEE Transactions on Systems, Man, and Cybernetics.
- Norman, D. A. (2013). The Design of Everyday Things. Basic Books.
Design thinking and technology adoption
- Brown, T. (2008). Design thinking. Harvard Business Review.
- Gartner Research. Hype cycle methodology.