Walk a Toyota store's used lot and count the badges. You will find Fords, Chevrolets, Hondas, and a few makes the store has no institutional reason to understand. That is not an anomaly, it is how franchise used-car retailing works everywhere. By Autonomi's count, somewhere between a quarter and a third of retail used deliveries at a typical franchise location are units the franchise OEM never built a single training program or incentive around.
Every one of those units was acquired by a person making a judgment call in an auction lane or at an appraisal desk, reasoning from instincts built stocking a completely different brand. The store's own answer to whether that was a good call already exists. It is in the sold log, and almost nobody has ever read it that way.
Disclosure: Floof Digital was the platform engineering partner on the Autonomi build, so we have a stake in the tooling described here. Discount the enthusiasm accordingly. The underlying defect is worth your attention whether or not you ever buy anything, because it is not really about cars.
The taxonomy came from the other department
Here is the actual failure, stated generally, because it will be familiar to anyone who has never set foot on a dealer lot.
The used side inherited its classification scheme from the new side, where that scheme is correct. Franchise brand is the organizing dimension of the new-car business. OEM certification, reconditioning standards, wholesale guidance, co-op reimbursement, historical comps: all of it is built around the primary brand, by design, and it works. Then the same lens gets pointed at the used lot, where a third of the inventory has nothing to do with the franchise, and it stops describing reality.
Critically, it does not stop rendering. The report still runs. It still produces a number. The franchise units get analyzed accurately and the rest get swept into a line item called "other," or flattened into an undifferentiated used-car total. Nothing errors. Nobody gets an alert that a third of the business has fallen out of the analysis, because from the system's point of view nothing went wrong.
That is the whole problem in one sentence. A classification that no longer fits does not fail loudly. It fails by quietly producing a category named after the absence of a category, and then everyone downstream treats that category as small.
Go find your own bucket called other
The dealer version is vivid, but this pattern is close to universal, and we have walked into it in nearly every engagement where somebody says "we have the data, we just cannot get an answer out of it."
Revenue by product line, where a third of it lands in "misc" or "professional services." Support volume by issue type, where the largest single category is "general inquiry." Marketing spend by channel, where "direct" and "unattributed" together outweigh every named source. Customers by segment, where the biggest segment is the one you invented to hold everyone who did not fit the other four. Headcount by function, where "operations" means eleven unrelated jobs.
In each case, the bucket is not a rounding category. It is the part of the business nobody has a lens for, and it is usually large precisely because it accumulated everything the original scheme could not describe. Size it once and the diagnosis is immediate. If more than about a tenth of anything that drives a decision lands in "other," you are not looking at a tidy remainder. You are looking at the section of your business that is managed on instinct, and its size is a measure of how far your taxonomy has drifted from your operations.
The three costs, and the one nobody measures
Autonomi's breakdown of what a gut-feel acquisition actually costs is worth borrowing wholesale, because the structure generalizes.
Acquisition price is visible the moment the deal is booked. Overpay for a make that turns slowly at this store and the gross is compressed before the unit ever hits the lot. Most stores track this. Far fewer trace it back to a pattern in acquisition decisions, because doing so means joining acquisition price against sold velocity by make, and that join does not fall out of a standard DMS export.
Days to turn is worse because it is a meter, not an event. Floor plan interest accrues on the outstanding balance whether the car moves or not. A unit aging on the off-brand row is not merely failing to produce gross, it is actively producing carrying cost every day, and the usual ending is a progressive series of price cuts followed by a wholesale exit at a loss. The decision that set that in motion happened on a Tuesday in a lane, on instinct, with no historical context available.
The trade you did not take is the one that matters most and gets measured never. Every unit the desk overpays for and then watches age is a unit that did not come in. Compound that across twelve months of acquisitions and it dwarfs the other two.
Notice the ranking. The cost that is easiest to see is the smallest, and the cost that compounds is invisible. That is not a coincidence and it is not specific to cars: measurement selects what gets managed, so the unmeasured cost is reliably the one allowed to grow. We made a version of this argument about production dashboards in Faster Than You Can Count. Same disease, different department.
Derive the classification. Do not configure it.
The best idea in the Autonomi piece is a small technical decision that deserves more attention than it will get.
To analyze off-brand units you first have to know which units are off-brand, which means knowing the store's franchise. The obvious implementation is a setting: a dropdown where the dealer picks their brand. Instead, AEGIS infers it from the store's own new-vehicle delivery history. Four hundred new Toyotas and three new Fords last year means the franchise is Toyota and every non-Toyota used delivery is off-brand. No dropdown, no field for the used desk to maintain.
That sounds like a minor engineering preference. It is not. A configured value is a promise a human has to keep. A derived value is a property of the data. Configuration rots the moment someone is on vacation during an acquisition, or a second franchise gets added, or the person who understood the field leaves. Every stale dropdown in every system you own was correct on the day it was set. Derivation cannot go stale, because it recomputes from the same evidence the analysis is already reading, and when the underlying reality changes the classification changes with it, without anyone remembering to do anything.
We keep arriving at this same principle from different directions. It is the identical instinct behind our piece on who actually owns the ad account: prefer the arrangement that is structurally true over the arrangement that depends on someone honoring it later.
The boring part is the actual work
Before any of that analysis can run, three unglamorous problems have to be solved, and they are where these projects genuinely die.
Every DMS exports differently, so column names and layouts vary and a log that served one purpose is often missing exactly the field a different question needs. The new-versus-used flag is sometimes explicit, sometimes inferable, sometimes simply absent. And the make field itself is entered by humans, so "CHEV", "Chev", "Chevrolet" and "chevy" are one brand wearing four labels, and any count that does not collapse them undercounts every brand with label variation, which is all of them.
That last one deserves a moment because it is the most common way an analysis lies to you while appearing to work. Split one brand across four spellings and it shows up as four small entries instead of one large one, none of them big enough to notice, all of them ranked below brands whose names happen to get typed consistently. The output looks fine. It is confidently, invisibly wrong.
People treat normalization as the tedious preamble before the real analysis. It is not the preamble. On this class of problem it is the analysis, and the model on top of it is comparatively trivial. The right posture on missing data is worth stealing too: derive a field from the evidence where that is honest, flag the derived values as derived, and fail the upload with the exact missing fields named when the file genuinely cannot support the question. Never quietly invent a value to keep the pipeline moving.
The playbook
DMAIC at the strategic layer, disciplined delivery underneath. Six moves, and none of them require buying data you do not already own.
Find your "other" and size it this week. Pick the three reports your decisions actually rest on and measure the residual bucket in each. The number itself is the finding, and it usually shocks people.
Canonicalize labels before you count anything. One mapping table from every observed spelling to one canonical value, applied at read time. Do this first, because every number you produce before it is wrong in a direction you cannot predict.
Derive classifications from behavior, not from a maintained field. If a fact can be inferred from what the business actually did, infer it. Reserve configuration for things that genuinely have no evidence behind them, and treat every remaining dropdown as a maintenance liability with a name attached.
Measure the decision, not just the outcome. Cohort results by the choice that produced them, which for a used desk means acquisition price against sold velocity by make. Outcome reporting tells you what happened. Decision reporting tells you what to do differently, and almost nobody builds the second one.
Trust your own history over the market's. Market data tells you what a make is worth to the median buyer in your region, and you are not the median. The same unit can turn in three weeks at your store and sit for two months at a competitor across town because the traffic profiles differ. Local signal beats broad signal on any decision that is actually local.
Put the judgment on a clock. A stocking posture is not a configuration value you set once. A make that sat all spring may move this fall because the market around it shifted. Any signal derived from history has to be recomputed on a schedule, or it becomes exactly the stale dropdown you just replaced.
Our position
What makes this problem worth writing about is that nothing is missing. The store is not short of data. It has recorded every single one of these decisions and their outcomes, faithfully, for years. It simply files a third of them under a heading that guarantees nobody will ever look, and then buys market-level subscriptions to answer a question its own records already answer better.
That is the shape of most analytics failures we get called into. Not an absent data set, a wrong lens: a classification inherited from a neighboring part of the business, correct where it came from, quietly useless where it landed, and never re-examined because the report keeps rendering. The fix is rarely a purchase. It is usually one derived column and a mapping table, which is unglamorous work with an embarrassing return.
Go look at whatever your business calls "other." Measure how big it is. If a third of your operation is filed under the absence of a category, that is not a reporting quirk. That is where the money is.