Watch a team estimate for twenty minutes and you can usually predict their next three retrospectives. The tells are all there. The senior dev who thinks out loud before anyone votes. The product owner glancing at the sprint capacity spreadsheet mid-discussion. The fifteen-minute standoff over whether something is a 5 or an 8, as if anyone will remember the difference in a month.
None of these habits look like problems in the moment. They look like collaboration, diligence, rigor. That's what makes them dangerous. Each one quietly bends the numbers away from what the team actually believes, and once the numbers stop meaning anything, everything built on top of them (forecasts, commitments, trust) starts to wobble too.
Here are the six I see most often, what's actually driving each one, and the counter-move that breaks it.
Anchoring: the architect speaks first
The story comes up on screen. Before anyone has voted, the most senior person in the room leans back and says, "Hmm, feels like a 5 to me." Then everyone votes, and what do you know: a wall of 5s, with maybe one brave 8 from the person who actually maintains that module. The discussion that 8 should have triggered never happens, because the room has already converged.
This isn't weakness or laziness. Anchoring is one of the most reliable effects in decision research: the first number spoken becomes the reference point everyone else adjusts from, and people adjust far less than they think they do. Seniority makes it worse. Disagreeing with the architect's casual guess costs social capital, and a junior dev who suspects the story is twice as big will quietly decide it's not worth the fight.
The counter-move is mechanical, which is why it works: silent, simultaneous reveal, every single time. Nobody muses, nobody floats a "rough gut feel," nobody gets a preview. Everyone commits to a number in private and all the cards flip at once. No exceptions, and especially not for the architect, because the architect is precisely the person whose stray comment flattens the vote. This is also the cheapest fix on this list. Simultaneous reveal is one click in any decent planning poker room; the hard part is enforcing the "no thinking out loud first" rule, and that's a five-second interruption from whoever is facilitating.
Estimate-to-fit: sizing backwards from the answer
You can hear this one happen. "Well, we've got about 30 points of capacity left, and we really need these four stories in... so this one's probably a 5, right?" The number wasn't estimated. It was reverse-engineered from the answer someone needed. The deadline version is uglier: the release date is fixed, the scope is fixed, so the estimates become whatever makes the arithmetic work, and everyone in the room knows it.
The incentive underneath is simple: the estimate is the only variable anyone feels allowed to touch. Scope has been promised, the date has been promised, so the number becomes the pressure valve. The problem is that estimating-to-fit doesn't change the work. The story that got squeezed into a 5 is still an 8's worth of effort. You haven't made the sprint fit; you've made the sprint lie, and the lie gets discovered around day seven.
Break it by strictly ordering the two conversations. Estimate first, with the capacity spreadsheet closed and the deadline off screen. Only then open the planning discussion and let it react to the estimates: cut scope, move the date, split the story, whatever the honest numbers force you to do. Estimation answers "how big is this?" Planning answers "what do we do about that?" The moment those two questions get asked simultaneously, the second one always wins.
The padding arms race
A dev privately thinks a story is a 5, calls it an 8, and feels clever. Her manager, who has seen this movie before, mentally discounts every estimate by a third and pushes back on all of them, including the honest ones. So next quarter the devs pad harder, because pushback is now guaranteed. The manager cuts deeper. Within a year, nobody in the building knows what any number actually means, and both sides are certain the other one started it.
Both sides are behaving rationally, which is why appeals to honesty don't fix it. Devs pad because unpadded estimates get treated as commitments and blown commitments get punished. Managers cut because they correctly suspect padding. It's a textbook trust spiral: each side's defense is the other side's evidence.
The way out is to stop hiding uncertainty inside the number and put it next to the number instead. A single figure like "8" forces the estimator to smuggle their doubt into the size, and forces the reader to guess how much doubt is in there. A range with a confidence signal ("probably 5, could be 13 if the legacy import code is as bad as we fear") smuggles nothing. The uncertainty is on the table where it can be discussed, investigated, or priced in, instead of silently inflating everything. Some teams run a quick confidence vote after the size vote; others just say the range out loud and write it on the ticket. It matters less how you express it than that padding stops being the only channel for expressing fear.
Velocity theater
The team's velocity was 34, then 38, then 42, and a slide somewhere celebrates the trend. Meanwhile nothing about the actual output has changed. Stories that were 3s last quarter are 5s now. Nobody sat down and decided to inflate; the drift happened one "eh, call it the bigger number" at a time, ever since velocity started appearing in a management dashboard next to other teams' numbers.
This is Goodhart's law doing exactly what it always does: when a measure becomes a target, people optimize the measure instead of the thing it was supposed to measure. Velocity is a fine internal signal and a terrible KPI. The instant it affects how a team is judged, the team will (consciously or not) protect it, and the cheapest way to protect a number the team itself generates is to redefine what a point means. Comparing velocity across teams is even more hollow, since points are relative units that only mean anything inside the team that calibrated them. I've written more about that relativity in what story points actually measure.
The counter-move requires a manager to give something up, which is why it's rare: keep velocity inside the team, full stop. Use it for what it's good at, which is turning your own history into a forecast (that's the whole argument of forecasting with velocity instead of promises), and never put two teams' velocities on the same chart. If leadership needs a health signal, give them cycle time, escaped defects, or delivered outcomes. Anything the team can't inflate by renaming its own units.
The precision trap
Twenty minutes into refinement, two developers are still litigating whether a story is a 5 or an 8. Both have made the same three arguments twice. The rest of the team has mentally left the building. Eventually someone splits the difference in spirit, everyone agrees to "5, but a big 5," and the meeting moves on having produced one estimate at a cost of roughly two hours of combined attention.
The bias underneath is the belief that estimation accuracy is a function of discussion time. It isn't, not past the first few minutes. The genuinely useful part of a 5-versus-8 disagreement (what does the 8 voter know?) surfaces almost immediately. Everything after that is precision theater over a difference that sits comfortably inside the noise floor of any estimate. Your actuals for "5-point stories" already vary by more than three points. You are debating decimal places on a measurement taken with a bathroom scale.
The fix has two parts. First, a default rule: when the room is split between adjacent values and the discussion has stopped producing new information, take the higher number and move on. The pessimist usually knows something, and if not, you've overestimated one story by one card. Second, remember why the deck has gaps in the first place. There's no 6 or 7 between 5 and 8 precisely so you can't split hairs there. The deck is telling you the distinction you're fighting over doesn't exist. Listen to it.
Estimating alone
The tech lead sizes the whole backlog on a quiet Friday afternoon. It's efficient, the numbers are internally consistent, and refinement meetings get short. Then the sprint starts, a mid-level dev picks up a "3" that touches code she's never seen, and it takes her a week. She never agreed to that 3. She never saw it before it landed in her queue. But it's her name next to the overrun.
This one usually comes from good intentions. The lead is fast, estimation meetings are expensive, and handing the team pre-sized stories feels like a favor. Sometimes it's rationalized as expertise: the lead knows the codebase best, so the lead's numbers are the most accurate. Except accuracy was never the main point. An estimate is partly a prediction and partly a commitment, and a commitment made on your behalf by someone else is neither. Solo estimates also encode exactly one person's knowledge and one person's blind spots. The dev who would have said "that module has no tests, I'm voting 8" was never in the room.
The counter-move is simply the definition of team estimation: whoever might do the work, votes. That's the entire point of the ritual. The spread of votes is where the hidden knowledge lives, the discussion is where it gets shared, and the final number is something the team can stand behind because the team produced it. If refinement is too expensive to run with the whole team, fix the meeting (fewer stories, better preparation, an async round of voting before the call) rather than removing the team from it.
Fix the one your last retro complained about
Six anti-patterns is too many to attack at once, and you don't have all six anyway. You have one or two, and your team already knows which. Go read your last retro board. "Estimates keep being wrong" next to a certain senior name is anchoring. "We committed to too much again" is estimate-to-fit. "Refinement takes forever" is the precision trap wearing a scheduling complaint as a disguise.
Pick that one. Make the counter-move a standing working agreement, write it where the team votes (ScrumMastr lets you pin notes to a room, or a sticky on the wall works fine), and give it three sprints before judging. Most of these fixes are procedurally tiny. Reveal simultaneously. Estimate before planning. Take the higher card. The habit is small; the honesty it buys back is not.