Great Martinis & Algorithms are DRY
Functional Programming Isn't Just for Academics — Part 23
Many years ago early in the Fall, I was in a bar. The only other patron was the bartender's classmate. They were taking a numerical methods class and they were scratching their heads over an assignment that involved (among other things) shuffling a deck of cards and dealing them over an arbitrary number of players in C. While they were debating using FFT (Really?!?) to perform the shuffle. I grabbed a pen and the cocktail napkin under my martini and wrote a function that returned the difference of two random numbers, slipped it to the bartender and told him he could quick sort his cards using stdlib and a pointer to my silly function. For the rest of the semester he got A's and I got very generous pours...
... 30 years later marketing asks for customer clustering. Reasonable enough. We have customers. We have orders. We can calculate recency, frequency, spend, promotion sensitivity, return behavior, channel preference, whatever matters to the business, hand the resulting vectors to a clustering algorithm and give marketing some segments... Then merchandising wants product clustering. Different Jira ticket. Different inputs. Different business owner. Products have attributes instead of demographics, attach rates instead of purchase histories, markdown sensitivity instead of lifetime value. So we build product clustering... Then somebody wants to group promotions by the customers and products for which they actually perform. Content wants something similar for campaigns and experiences. Marketplace wants it for sellers. Operations wants it for orders, fulfillment patterns and returns.
Somewhere along the way we have accidentally written five versions of substantially the same calculation. The usual response is to extract common code once the duplication becomes embarrassing. I think that starts one question too late. The useful question was available before the first line of customer clustering was written: What part of the clustering algorithm actually knows what a customer is?
