Expert Advice · Training · 14 min read
Classical vs operant conditioning in dog training
Classical pairs events. Operant changes what the dog does next. “Balanced” is a tool list, not a third law. Run one loop you can name.
Classical conditioning pairs events so the dog predicts what is coming. Operant conditioning changes what the dog does next because a consequence followed the behavior. In a kitchen both run at once — the bowl click is classical, the sit that produces the bowl is operant. “Balanced training” is a shopping list of tools, not a third learning theory. If you cannot name the loop, you are wishing.
I write this as a trainer, not as a lecturer in a lab coat. Detection dogs and pet dogs sit in the same processes. The criteria change. The laws do not.
What do CS, US, CR actually mean?
Classical notation is four letters you can hang on a kettle.
- US — unconditioned stimulus. The thing that already works without training. Food in the mouth. A sudden bang that startles.
- UR — unconditioned response. The unlearned reaction. Salivate. Flinch.
- CS — conditioned stimulus. A previously boring event that, after pairings, predicts the US. The kettle click. The lead jingle. Your hand on the treat pouch.
- CR — conditioned response. What the dog now does when the CS happens. Orient to the bowl cupboard. Light up at the lead.
The dog does not “decide” to salivate when the kettle sounds like dinner. The pairing did that. You can use this on purpose: the crate door, the mat, the clip of the lead become CS for something decent instead of CS for isolation or a fight. You can also poison a CS without noticing. If the click of the lead always means a chaotic pavement, the jingle is no longer neutral. That is still classical. It is just a pairing you did not vote for.

What do Sd, R, Sr actually mean?
Operant notation is the loop you run on purpose.
- Sd — discriminative stimulus. The cue or the picture that sets the occasion. You standing at the door. The word “sit.” The sight of a bowl in your hands.
- R — response. The behavior you can see. Hip on the floor. Four paws still. A nose touch.
- Sr — reinforcing stimulus. What follows and makes R more likely next time. A piece of the ration (Sr+). Or the removal of something the dog wants gone (Sr−) — the lead slackening, you stepping away from a tight space.
Punishment exists in the same grammar. Something aversive is added, or something the dog wanted is taken away, and the behavior gets less likely. I am not going to pretend the process is imaginary. I am going to say that on this lab the default loop is reinforcement you can see, in a setup the dog can win. If you add an aversive, you still have to name it as punishment or negative reinforcement. Calling the collar “balanced” does not change the letters.
Public overviews of the operant skeleton — Britannica’s entry on operant conditioning is a clean one — describe the same three-term contingency trainers use on a kitchen floor. Read the original. This page is a working translation, not a reprint.
How does this look in an ordinary kitchen?
Classical, kettle. You boil water every morning. After a week the click of the kettle is CS for movement toward the worktop, because that click has predicted food-adjacent bustle (US). Nobody taught a cue. The pairing taught a prediction.
Classical, lead. The lead lives on a hook. If every walk is a rehearsal of lunging, the hook-and-jingle becomes CS for arousal before you have said a word. If every clip is followed by a scatter of food on the floor and a boring first twenty metres, the same jingle becomes CS for “we start quiet.” Same object. Different pairing.
Operant, greeting. You walk in. Sd is you-at-the-door. If four paws on the floor (R) produce a greeting and a piece of the ration (Sr), floor-standing gets likelier. If jumping produces hands and a voice, jumping gets likelier. You are already running a loop. The only question is which R you are paying.
Operant, bowl. You pick up the bowl (Sd). Sit (R) produces the bowl on the floor (Sr). Stand and shuffle and the bowl waits. That is not manners as morality. It is a contingency the dog can read.
Most minutes contain both. The bowl pick-up is a CS for food (classical) and an Sd for sit (operant). Trainers who pick a fight between the two words are usually trying to win an internet argument. On the floor you need both names so you know what to change.
What is an extinction burst, and why do people call it spite?
Extinction is what happens when a previously paid behavior stops paying. The first thing you often see is not silence. It is a bigger version of the old behavior — faster jumping, louder barking, a sit that turns into a paw and a whine. That spike is the extinction burst. If you give in during the spike, you have taught the louder version. If you hold the new contingency and the dog can still win another way, the old R dies.
A household example: you used to pay jumping with hello. You stop. Night two the dog jumps higher and vocalises. That is the burst, not a moral collapse. You become furniture, you pay four paws, you do not narrate. If the dog cannot win any greeting at all, you have not run extinction. You have run frustration with no legal door. Always give a paid alternative.
Do not run extinction on a behavior that is keeping the dog safe, and do not run it on panic. Panic is not an operant you wait out. See when to call the vet if the body is the loud fact.
Why is “balanced” not a learning theory?
Learning theory names processes. Pairing. Reinforcement. Punishment. Extinction. Generalisation. “Balanced” names a toolkit — food plus a leash pop, or a click plus an e-collar, or “I use everything.” That is a purchasing decision. It is not a fourth process.
I have sat in rooms where a handler said they were balanced because they used treats and corrections. The dog was still in Sd–R–Sr, with an aversive Sr when the handler disliked the R. The letters did not change. The welfare did. If you use an aversive, say so in process words. Then ask whether the same criterion could be taught by changing the setup and paying a legal behavior. On this lab that is the default question.
I am not going to run a brand war. I am going to refuse the idea that a marketing adjective replaces Pavlov and Skinner. If a method cannot tell you what the CS is, or what the Sr is, it is not a method. It is a vibe.
How do you run one operant loop tonight?
Name one behavior. Four paws on the floor is enough. Set a boring kitchen. Wait or lure a cheap version. Mark the instant the criterion is true. Pay from the ration. Reset. Eight to twelve wins. Stop.
That is the HowTo on this page. The markers and criteria article is the next hinge: how often the dog should win, and when you are allowed to make the picture harder. Food logistics — so the pouch is not a second dinner — live on food as a reinforcer.
If you want a skill ladder rather than a loop, go to skills. If the mouth or the crate is the actual Tuesday problem, stay in puppy. Theory that does not change Thursday evening does not belong on this site.
Sources
Encyclopaedia Britannica, “Operant conditioning,” public overview of the three-term contingency. Paraphrased as a learning-theory skeleton, not copied as a kitchen protocol.
Questions on this page
What is the difference between classical and operant conditioning?
Classical pairs a neutral event with something that already matters, so the dog predicts it. Operant changes what the dog does next because a consequence followed the behavior. Most kitchen minutes contain both.
What do CS, US, CR and Sd, R, Sr actually mean?
In classical, US is the thing that already works (food in the mouth), UR is the unlearned response, CS is the signal that starts to predict it, CR is the learned response to that signal. In operant, Sd is the cue or setup, R is the behavior, Sr is the reinforcer that follows.
What is an extinction burst?
When a previously paid behavior suddenly stops paying, the behavior often gets bigger before it dies. People call this spite. It is the burst. If you pay the burst, you have taught a louder version.
Is “balanced training” a learning theory?
No. Learning theory names processes — pairing, reinforcement, punishment, extinction. Balanced is a marketing cluster of tools. Using an aversive is still operant punishment or negative reinforcement.
Do I need a clicker to run an operant loop?
No. You need a named behavior, a setup the dog can win, a mark if you want the timing clean, and a pay. A quiet “yes” is enough. The clicker page is a separate protocol.