Why LLM subagents, not regex, for deduplicating legacy item codes
You bring a 20-year-old product master into Odoo, and you find the same item filed under three or four codes. My migration had roughly 25% of legacy codes as duplicates. The natural reflex is to write a regex. I did. It caught the easy ones and missed the ones that cost real time downstream. The approach that worked was purpose-built LLM subagents running category by category. Why do regex rules fall apart on real item codes? Regex is exact. It matches strings you can describe precisely. Legacy item codes are not precise. They are the residue of years of manual entry: truncations, spelling variants, swapped words, and category words that drifted out of a line name. A rule that matches “Mandarins 5kg” does not match “Mand 5kg” or “5kg Mandarin” or “Mandarin (5kg)” depending on which operator typed it and which year. Each exception needs a new rule, and each new rule needs a new test. You are not fighting one dirty field. You are fighting a thousand small human decisions. ...