WRITTEN BY: HARRIE PHILLIPS
PGCertClinEd, BAdVocEd (VocEd&Trng), DipVN (Surgical, ECC), DipBus, DipTAE (Development & Design), TAA
PUBLISHED: 12 June 2026
Horses dislike surprises. Anyone who has been around horses for even a split second knows this. The plastic bag in the hedge that wasn’t there yesterday, the new horse in the next paddock, the wheelie bin that has somehow become threatening overnight, all of these produce the same response: head up, eyes wide, body braced. And if we are really unlucky, some fancy footwork, a spin and speed even Usain Bolt would be proud of. As prey animals, horses are evolved to treat the unexpected as dangerous until proven otherwise, and we spend a lot of our training lives trying to convince them that the world is more boring than it looks. So why am I saying “reward” and “surprise” in the same sentence as “horse”?
Because there is another kind of surprise, the kind that happens inside the horse’s brain when a reward arrives that they were not quite expecting. This kind of surprise does not necessarily produce alarm. Instead, it can generate a learning signal when an outcome differs from what the brain expected. It is one of the important neural mechanisms through which mammals update expectations about cues, actions and outcomes, and it has useful implications for how we train horses.
This article is about that second kind of surprise. Where it comes from, why it matters, what the most current research from primates, mice and horses tells us about how it works. And I am hoping I can translate that into something understandable, something useful for the way you train, ride and handle your horse from tomorrow onwards.
What does dopamine actually do in your horse’s brain?
Dopamine has a reputation it does not quite deserve. Most people, if they know anything about it, know it as the “pleasure chemical”, the brain’s reward signal, the molecule that makes things feel good. This is not exactly wrong but it is misleading enough to obscure what dopamine is really doing in any mammalian brain, including your horse’s.
One major function of phasic dopamine signalling is to encode a reward prediction error1, 2: the difference between what the brain expected and what actually happened. When an unexpectedly good outcome occurs, dopamine neurons in the midbrain show a brief burst of activity. Think about how happy you’d be if you finally won lotto after many years of buying tickets. When the outcome is fully predicted, there is little or no prediction-error response at the moment the reward arrives. When an expected reward does not occur, dopamine activity can briefly fall below baseline. Disappointment, anyone?
Importantly, that does not mean a predicted reward has stopped being rewarding. As learning develops, the dopamine response can shift from the reward itself to the cue or event that predicts it. The brain has learned the relationship. Dopamine, in this context, is therefore not simply encoding ‘yum’. It is helping signal whether an outcome was better than expected, worse than expected, or about what the brain predicted.
Applying this framework directly to horses still requires some extrapolation. Direct recordings of dopamine prediction-error signals in behaving horses are not currently available, so equine researchers infer the likely role of these systems from behaviour, physiology, equine neurobiology and comparative mammalian neuroscience3, 4.
Recent equine research does, however, fit well with the broader framework. Work at Nottingham Trent University has shown that physiological arousal is associated with performance on learning tasks5. McBride and colleagues have reviewed how the equine basal ganglia and dopaminergic systems are likely to contribute to learning6, while Henshall and colleagues have discussed equine learning, including aversive learning, using prediction-error mechanisms as part of the explanatory framework7.
The cautious conclusion is not that horses have been shown to produce exactly the same dopamine signals recorded directly from laboratory rodents and primates. Rather, the wider mammalian evidence provides a biologically plausible and useful framework for interpreting equine learning, while direct equine neurophysiological evidence remains more limited.
Dopamine is not simply a ‘pleasure signal’. In reward learning, prediction-error signalling helps the brain detect when an outcome is better or worse than expected and update its predictions accordingly.
Why does this matter for everyday horse training?
So, if you’ve stuck with me after that last bit, well done. Here is the part that should change how you think about training.
A reward becoming predictable does not mean that it stops being a reward or stops reinforcing behaviour. In early learning, clear, prompt and reliable reinforcement is usually exactly what we want because it helps the horse identify which response produced the outcome. As the horse learns that relationship, the prediction-error response at the moment the expected reward arrives becomes smaller because the outcome is no longer surprising2, 8. That is evidence that the horse has learned to predict the outcome, not evidence that reinforcement has stopped working.
This is where two related but different ideas need to be separated. Prediction-error research tells us that outcomes that are better or worse than expected provide information that updates expectations9. Reinforcement-schedule research tells us that, once a behaviour is well established, intermittent or variable reinforcement can make that behaviour more persistent and resistant to extinction than continuous reinforcement10. Those findings do not mean that variable reinforcement is inherently better for teaching a new response.
During acquisition, the horse needs a clear contingency: this response produces this consequence. Reinforcing the correct response promptly and reliably helps establish that relationship. Once the behaviour is well learned, reinforcement can sometimes be thinned or varied. An occasional unexpectedly valuable outcome can also create a positive prediction error and may help highlight an especially good response.
For horse training, the practical implication is therefore more nuanced than ‘make rewards unpredictable’. Keep the cue, the response criterion and the relationship between behaviour and consequence clear. Predictability in the training rule is a strength. Surprise can sometimes be useful within that predictable framework.
What’s the trap most people fall into?
Now, before anyone reading this starts varying their cues or their timing in pursuit of more dopamine, let me be very clear, because this is the part that goes badly wrong if it is misunderstood.
If you choose to introduce variability once a behaviour is well established, it belongs in the reward rather than in the cue or the timing of the consequence.
The cue you give your horse for forward should mean the same thing every time. The cue for halt should mean the same thing every time. The release of pressure or the marker for a positive reinforcer should occur promptly and consistently when the correct response begins. Where a marker is used, it can bridge the short delay until the primary reinforcer is delivered. These three pillars of clarity of cue, timing, and consistency are the foundation on which all training rests, and they do not change because I am now banging on about varying something.
Once a behaviour is well established, what can vary is the type, value or frequency of any additional reward that follows the release or marker. Sometimes the established reinforcer may be sufficient. Sometimes you may add a wither scratch, food, rest, enthusiastic praise or another outcome that the individual horse values. Variation can be used to thin reinforcement appropriately, maintain engagement and occasionally make an especially good response produce a better-than-expected outcome.
A 2025 narrative review by Bradshaw-Wiley, Henshall, McLean and Freire examined combined positive and negative reinforcement in equine training11. The authors identified combined reinforcement as a promising area for further investigation, but also highlighted how limited the direct evidence remains and how inconsistently reinforcement methods have sometimes been described in equine research. The review therefore does not establish that combining pressure release with a positive reinforcer is inherently more powerful than pressure release alone. Instead, it highlights the need for better research into how combined approaches affect learning and horse welfare.
If you take one practical idea from this section, take this: keep your cues and timing consistent. Once a behaviour is established, the rewards around that behaviour can sometimes vary. The cues are the words. Timing and contingency are the grammar. Rewards are the feedback. Keep the words and grammar clear, while the feedback can sometimes vary.
Keep your cues consistent. Rewards can vary once the behaviour is established.
What does Janet Jones add to this?
Dr Janet Jones is a cognitive neuroscientist who applies brain research to horse training, and she can explain these concepts much better than me. If I have even piqued a little of your interest whilst reading this, her books definitely should be on your bookshelf.
In her 2020 book Horse Brain, Human Brain and in her widely-read Psychology Today column, Jones talks about how food rewards should be used rarely, not routinely12. Jones applies general reward-prediction error research to argue for relatively infrequent food rewards. That is an interesting training hypothesis, but it should be presented as her interpretation rather than as an established conclusion of equine neuroscience. Peer-reviewed horse studies have not shown that frequent, correctly contingent food reinforcement “wastes” training power because it becomes predictable.13.
Her recommended approach is to lean on non-edible rewards (scratches, praise, rest, the cessation of work) as the everyday currency of training, and to save food for rarer, higher-value moments where she argues that the surprise may make the outcome especially salient. A treat that comes out of nowhere after a particularly good attempt at something difficult, in her framing, is a more powerful learning event than a treat that is dispensed reliably after every repetition of an established behaviour.
If you follow Jones’s approach, it also means you would avoid ‘wasting’ a treat by giving your horse one for nothing. I get it, they’re cute and all, but are you giving them the treat for them, or for you? That’s what I thought. Save it for something much more meaningful.
This is a stronger position than the consensus within much of the equine positive-reinforcement community, where regular food reinforcement is widely used. Jones takes the reward-prediction-error literature further and uses it to argue for relatively infrequent food rewards. That is an interesting practical interpretation, but it should not be confused with evidence that predictable food ceases to function as a reinforcer. Regular, contingent food reinforcement remains a well-established training approach.
What is well established is that outcomes that are better or worse than expected generate prediction errors and can update learning. A trainer might make practical use of that through an occasional high-value ‘bonus’ reward or by thinning a reinforcement schedule once a behaviour is well established.
Does the same principle apply to pressure-release?
Most reward-prediction-error research has focused on appetitive rewards, but dopamine is also involved in learning about aversive events, escape and relief. The underlying circuits and signals are not identical to appetitive learning, so it is better not to describe positive and negative reinforcement as neurologically interchangeable. The important point is that prediction and prediction error can also contribute to learning when an animal discovers how its behaviour can terminate or avoid an aversive event.
Prediction errors help update learning. Clear contingencies tell the animal how its behaviour changes the outcome.
The neuroscience
Until recently, the role of dopamine in negative reinforcement learning was much less well-studied than its role in positive reinforcement. That gap has been closing rapidly, and what the science is showing is actually really interesting. Trust me. You must, because you are still reading.
In 2021, a group of researchers led by Zhijun Diao recorded directly from substantia nigra dopamine neurons in mice learning to escape mild foot shocks14. They found that the dopamine system was bidirectionally engaged throughout negative reinforcement learning, not silent as some earlier theories had assumed. In the early phase of learning, when the mice had not yet figured out how to escape, the aversive stimulus suppressed dopamine activity, as you would expect. Like when we accidentally touch the electric fence, it’s not exactly fun. But once the mice had learned the escape response, the same aversive stimulus became associated with a brief increase in dopamine activity once the escape response was well learned, consistent with the stimulus having acquired a different predictive significance. Kind of like how when you let go of the fence, you know the shock will stop.
A 2025 study by Cheng and colleagues extended this work substantially15. Using more sensitive techniques, they tracked exactly how the dopamine signal shifts across the stages of negative reinforcement learning, and found a beautifully clean pattern:
- First, before the animal can escape, the dopamine response tracks the aversive stimulus itself.
- Then, as the animal begins to learn that the aversive stimulus reliably ends, the dopamine response shifts to track the termination of the stimulus. This shift is consistent with learning about stimulus termination and relief, although negative and positive reinforcement are not neurologically identical.
- Finally, once the animal has fully learned to escape on cue, the dopamine response shifts again, this time to track the onset of the aversive stimulus, because that onset now reliably predicts the relief that the animal can produce by performing the escape response.
These studies show that dopamine can participate in negative reinforcement learning and in the prediction of relief. They do not mean that release from pressure should be made unpredictable. Quite the opposite. In horse training, the animal needs control over the contingency: when the horse gives the correct response, the pressure should reliably and promptly reduce or cease.
Henshall and colleagues have discussed prediction-error processes as a useful framework for understanding equine learning, including aversive learning7, and reviews of equine neurophysiology identify dopaminergic systems as relevant to learning6. The defensible conclusion is therefore that negative reinforcement involves prediction and learning about relief, not that it stops teaching when relief becomes predictable.
Your horse is ALWAYS learning.
Something I find myself saying to my fellow horse peeps all the time is that every interaction with your horse is training, whether you intend it to be or not. Prediction error is one part of that much bigger learning picture. The horse is continually learning which cues predict which events, which responses change outcomes, and which behaviours are worth repeating.
As trainers, our job is not to keep the horse’s brain surprised all the time. It is to make the relevant contingencies clear. The horse should be able to predict what a cue means and what response will produce release or reinforcement. Within that predictable framework, unexpectedly good outcomes can sometimes provide additional information and help strengthen or refine learning.
So the practical question is ‘Was the cue clear? Was the response criterion appropriate? Did the consequence follow the behaviour I wanted? And was the horse in a physical and emotional state that allowed useful learning?’
A note on what makes this all possible
Everything in this article also depends on the horse’s physical and emotional state. A frightened, flooded, exhausted, painful or chronically stressed horse does not stop learning, but stress can change what is learned and can impair the acquisition or flexible expression of the response we are trying to teach.
Stress can alter attention, memory and behavioural flexibility16. In equine research, horses exposed to an uncontrollable stressor learned a negatively reinforced task less efficiently than horses given short ridden exercise, while the stressed and inactive groups did not differ significantly17. Higher cortisol concentrations during learning were also associated with requiring more trials to reach the learning criterion. The authors appropriately concluded that uncontrollable stress may impair learning, while short, moderate exercise may facilitate it.
Horses with lower physiological arousal have also performed better on some learning tasks5. That does not mean an anxious horse is incapable of learning. In fact, fear and avoidance learning can be very powerful. It means that high or uncontrollable arousal can interfere with the particular learning, behavioural flexibility and performance we are trying to achieve.
If your horse is struggling to learn, do not assume that the answer is a more surprising reward. First consider pain, fatigue, fear, arousal, environmental difficulty and whether the task, cue and reinforcement contingencies are clear.
You can read about stress, arousal, and why pushing through almost never works here.
References
- Schultz, W. (2016). Dopamine reward prediction-error signalling: a two-component response. Nature Reviews Neuroscience, 17(3), 183–195. https://doi.org/10.1038/nrn.2015.26
- Schultz, W. (2024). A dopamine mechanism for reward maximization. Proceedings of the National Academy of Sciences USA, 121(20), e2316658121. https://doi.org/10.1073/pnas.2316658121
- Schultz, W. (1998). Predictive reward signal of dopamine neurons. Journal of Neurophysiology, 80(1), 1–27. https://doi.org/10.1152/jn.1998.80.1.1
- Mirenowicz, J., & Schultz, W. (1994). Importance of unpredictability for reward responses in primate dopamine neurons. Journal of Neurophysiology, 72(2), 1024–1027. https://doi.org/10.1152/jn.1994.72.2.1024
- Evans, L., Cameron-Whytock, H., & Ijichi, C. (2024). Eye understand: Physiological measures as novel predictors of adaptive learning in horses. Applied Animal Behaviour Science, 271, 106152. https://doi.org/10.1016/j.applanim.2023.106152
- McBride, S.D., Parker, M.O., Roberts, K., & Hemmings, A. (2017). Applied neurophysiology of the horse; implications for training, husbandry and welfare. Applied Animal Behaviour Science, 190, 90–101. https://doi.org/10.1016/j.applanim.2017.02.014
- Henshall, C., Randle, H., Francis, N., & Freire, R. (2022). Habit formation and the effect of repeated stress exposures on cognitive flexibility learning in horses. Animals, 12(20), 2818. https://doi.org/10.3390/ani12202818
- Bech, P., Crochet, S., Dard, R., Ghaderi, P., Liu, Y., Malekzadeh, M., Petersen, C.C.H., Pulin, M., Renard, A., & Sourmpis, C. (2023). Striatal dopamine signals and reward learning. Function, 4(6), zqad056. https://doi.org/10.1093/function/zqad056
- Fiorillo, C.D., Tobler, P.N., & Schultz, W. (2003). Discrete coding of reward probability and uncertainty by dopamine neurons. Science, 299(5614), 1898–1902. https://doi.org/10.1126/science.1077349
- Ferster, C.B., & Skinner, B.F. (1957). Schedules of Reinforcement. Appleton-Century-Crofts.
- Bradshaw-Wiley, E., Henshall, C., McLean, A., & Freire, R. (2025). Combined reinforcement: Poisoned cue or a panacea for modern equine training? A narrative review. Applied Animal Behaviour Science, 291, 106718. https://doi.org/10.1016/j.applanim.2025.106718
- Jones, J.L. (2020). Horse Brain, Human Brain: The Neuroscience of Horsemanship. Trafalgar Square Books. ISBN 978-1-57076-948-1. View at publisher
- Jones, J.L. (2022). New Ways of Using Food to Train Horses. Psychology Today, “Horse Brain, Human Brain” column, February 2022. www.psychologytoday.com/au/blog/horse-brain-human-brain
- Diao, Z., Yao, L., Cheng, Q., Wu, M., Di, Y., Qian, Z., Wei, C., Liu, Y., Tian, Y., & Ren, W. (2021). Involvement of midbrain dopamine neuron activity in negative reinforcement learning in mice. Molecular Neurobiology, 58(11), 5667–5681. https://doi.org/10.1007/s12035-021-02515-6
- Cheng, Q., Liu, W., Yao, L., Xu, S., Wei, C., Zheng, Q., Wu, M., Han, J., Liu, Z., Ren, W., & Sun, Z. (2025). Dynamic changes of dopamine neuron activity and plasticity at different stages of negative reinforcement learning. Proceedings of the National Academy of Sciences USA, 122(45), e2509072122. https://doi.org/10.1073/pnas.2509072122
- Mendl, M. (1999). Performing under pressure: stress and cognitive function. Applied Animal Behaviour Science, 65(3), 221–244. https://doi.org/10.1016/S0168-1591(99)00088-X
- Henshall, C., Randle, H., Francis, N., & Freire, R. (2022). The effect of stress and exercise on the learning performance of horses. Scientific Reports, 12, 1918. https://doi.org/10.1038/s41598-021-03582-4


Every due care has been taken to ensure the information herein is based on sources Veterinary Nurse Solutions believes to be reliable, but is not guaranteed by us and does not purport to be complete or error-free. As such, we do not warrant, endorse or guarantee the completeness, accuracy, and integrity of the information. You must evaluate, and bear all risks associated with, the use of any information provided hereunder, including any reliance on the accuracy, completeness, safety or usefulness of such information. As part of our quality control of information contained within this document, it has been peer-reviewed by qualified animal care professionals.
Veterinary Nurse Solutions acknowledges that there is more than one way to carry out many of the tasks described within this website, and techniques omitted are not necessarily incorrect.