An interesting aspect of OR and business analytics practice is the chance to observe the impact of a decision support system (DSS) on paying customers who deploy it in real-life situations to improve business decisions. A good DSS tends to make users more innovative - they feel more secure that they have 'science' and automation working for them, and try to maximize benefits and progress toward becoming super-users.
A robust DSS should anticipate and react practically to a wide range of user inputs. In particular, if the DSS is based on a historical data-driven analytical model, inputs that are within the historical range would produce outputs roughly based on 'interpolations' - more reliable. On the other hand, an input not seen in prior history is (again, roughly speaking) equivalent to operating a heavy machine outside the functional limits, and may produce extrapolations that may make analytical sense, but little business sense. The latter input scenario makes the DSS look 'silly' because the business-savvy user would have provided a far better manual answer for such new inputs, because unlike the DSS that is just a computer program, they 'know'.
A bad DSS is one that is not trusted by users. There's more to it that good analytics. They may turn off the analytical portions and just use the automating features. Usually, it is equally if not more important to get the non-analytical portions right to ensure that customers are always getting some value for money. Then there's my pet 'timely failure theory' gleaned the hard way. If "infeasible" is an utterly unavoidable option in a DSS output for what ever reason, then make sure that such failures are immediate. Nothing irritates twitter-era humans more than having to wait 20 minutes to find out that nothing is going to happen.
In the ongoing cricket world-cup, the new umpire decision review system (UDRS) that predicts 'outs' using the Hawk-eye ball-tracking technology is turning out to be an example of a good DSS. The machine-free accuracy rate of umpires was 92%, which jumps to 98% with DSS-assisted decisions. If the DSS determines that the predicted result is marginal, the original decision of the umpire stands.
In the past, if the users (umpires) had a slight doubt, they would err on the side of caution. In cricket, that would amount to a 'not-out' ruling in favor of the batsman. That behavior appears to have changed a bit. Now, the decisions are slightly more 'aggressive' i.e., in favor of what the DSS decision is likely to be if a TV-referral were to be made. If the DSS is a more accurate predictor (which it appears to be), then this is likely to be a positive change in user-behavior, on average. A big chunk of the 6% error reduction may well have come from a reduction in the number of 'false not-outs'. Given that cricket today is skewed in the favor of batsmen, this is a welcome development.
On the other hand, cricket administrators have to work hard on defining a simple and effective operating range for the DSS. Otherwise, detractors (batsmen!) will use this as an excuse to get the DSS banished. Similarly, the time taken by the incumbent process to process and return a response to an input review is far too long . A batsman who is going to ruled 'out' by a damn machine after playing an idiotic shot doesn't want the global television audience to see replay after embarrassing replay of his moment of madness.
Sunday, March 20, 2011
Sunday, March 13, 2011
Analytics and Cricket - VI: Improving the Cricket World-cup Schedule
Introduction
The second-most watched single-sport tournament on the planet (after the soccer world cup) is underway in the subcontinent. The ICC cricket world cup is held once every 4 years and the latest edition is being jointly hosted by India, Sri Lanka, and Bangladesh. The original list also included Pakistan, but a militant attack on the Sri-Lankan national cricket team's bus there two years ago led to safety concerns. Fueled by the Indian economy, cricket is now a multi-billion dollar sport and the creation of a professional city-franchise based sports league in India three years ago is the latest example of this phenomenon. For some reason, cricket has always had an intimate connection with Operations Research (search for cricket in the archive). Today's post presents an overview of the potential benefits of using OR-enhanced scheduling in cricket.
Motivation
The 2011 edition stretches over six weeks, and considering the fact that the number of teams who qualified for the finals is only 14, the duration of the tournament appears to be far too long, and has come in for considerable criticism. For example, the FIFA soccer world cup with more than twice the number of teams in the fray was completed in one month. There have been quite a few instances where the idling time for teams between any two games has been close to one week. Given this (and to minimize the number of boring mismatches), a recommendation for the 2015 world cup has been to reduce the number of teams from 14 to 10.
The 2011 WC format
The 2011 tournament schedule can be found here. We provide a quick overview of the format and 'constraints' here. The teams have been divided into two groups of 7. Each team in a group plays every other team in its own group. The top 4 teams from each group after the league phase enter the knock-out stage that includes 4 quarter-finals, 2 semis, and a grand finale. Each game lasts up to 8 hours, and is typically played in a day-night format, starting around 2pm and ending by 10pm. The majority of the revenue comes from TV ads. The media rights for this edition were sold for $2 Billion to ESPN. League matches involving host nations (and India in particular) garner huge ratings. Marquee match-ups are typically scheduled during weekends. Given the magnitude of revenue at stake, one can guess that even a 1% improvement to the schedule in terms of increased viewership and reduced player fatigue (thereby resulting in better match quality for fans) can add a lot of value.
A sample of hard and soft constraints
A minimum gap of two days between successive matches is a must to ensure adequate time for rest and travel. Minimizing total travel cost and idle time appears to be desirable. Host nations play their games on home turf to the extent possible. Day games (9am start) are also possible, if the morning fog/mist/dew factors are not overwhelming. On the other hand, locations with consistent dew problems at night are better suited for day matches to ensure a more fair contest. Reserve days are required for games that are washed-out due to inclement weather during the knock-out stage of the tournament, and any feasible schedule must take that into account. Of course, the OR-driven Duckworth-Lewis rules for rain-interrupted games are used to maximize the chances of getting a fair contest under the circumstances.
The Big Picture
The global cricket season is itself pretty packed and the overall schedule has come in for much criticism due to player burn out as well as overselling the game. As a cricket fan, it is pretty obvious that the status-quo is so dismal that better scheduling can maximize long-term revenue while also minimizing burn-out. Overselling is a major issue, not only because it kills the golden goose (long-term fan involvement in the game), but the number of inconsequential games being played has lead to match-fixing and spot-fixing (similar to point-shaving in basketball). The money involved in illegal betting is mind-blowing, and many suspect that it already is or could become another source of income for terrorist groups operating in the Af-Pak area. One does not expect OR to help resolve all these issues, but it can certainly be applied to some of the key ones it is well suited for, and the cascading positive effects can make a difference to the overall situation.
Potential OR Approaches
The Traveling Tournament Problem (TTP) popularized by Dr. Mike Trick at CMU appears to be a good starting point for improving the schedule for future cricket tournaments, including the world cup and the IPL, and ultimately, the global cricket season itself. Exact or heuristic approaches that combine constraint programming with MIP models appear to be well-suited to such problems given the complex and 'idiosyncratic nature' of some the scheduling constraints that tend to be imposed in cricket, as well the value added by even small scheduling improvements.
Looking Forward: The 2015 World Cup and Beyond
The next edition with be the "Anzac" world cup, hosted jointly by Australia and New Zealand. The distances between some cities that are likely to host some of the matches can be enormous (e.g. Perth and Sydney are more than 4000 miles apart) or relatively tiny (e.g. intra-NZ games, or Australian east-coast games (Melbourne-Adelaide is less than 800 miles), and the value that can be obtained by adopting optimized schedules can be significant.
Conclusion
I hope this post presents a reasonable high-level overview of cricket tournament schedules and motivates interested OR'ers to further investigate this problem. As a cricket tragic and OR professional, I would be happy to contribute toward any such effort.
(To be submitted as an entry to the March-2011 Informs Blog challenge: OR and Sports)
The second-most watched single-sport tournament on the planet (after the soccer world cup) is underway in the subcontinent. The ICC cricket world cup is held once every 4 years and the latest edition is being jointly hosted by India, Sri Lanka, and Bangladesh. The original list also included Pakistan, but a militant attack on the Sri-Lankan national cricket team's bus there two years ago led to safety concerns. Fueled by the Indian economy, cricket is now a multi-billion dollar sport and the creation of a professional city-franchise based sports league in India three years ago is the latest example of this phenomenon. For some reason, cricket has always had an intimate connection with Operations Research (search for cricket in the archive). Today's post presents an overview of the potential benefits of using OR-enhanced scheduling in cricket.
Motivation
The 2011 edition stretches over six weeks, and considering the fact that the number of teams who qualified for the finals is only 14, the duration of the tournament appears to be far too long, and has come in for considerable criticism. For example, the FIFA soccer world cup with more than twice the number of teams in the fray was completed in one month. There have been quite a few instances where the idling time for teams between any two games has been close to one week. Given this (and to minimize the number of boring mismatches), a recommendation for the 2015 world cup has been to reduce the number of teams from 14 to 10.
The 2011 WC format
The 2011 tournament schedule can be found here. We provide a quick overview of the format and 'constraints' here. The teams have been divided into two groups of 7. Each team in a group plays every other team in its own group. The top 4 teams from each group after the league phase enter the knock-out stage that includes 4 quarter-finals, 2 semis, and a grand finale. Each game lasts up to 8 hours, and is typically played in a day-night format, starting around 2pm and ending by 10pm. The majority of the revenue comes from TV ads. The media rights for this edition were sold for $2 Billion to ESPN. League matches involving host nations (and India in particular) garner huge ratings. Marquee match-ups are typically scheduled during weekends. Given the magnitude of revenue at stake, one can guess that even a 1% improvement to the schedule in terms of increased viewership and reduced player fatigue (thereby resulting in better match quality for fans) can add a lot of value.
A sample of hard and soft constraints
A minimum gap of two days between successive matches is a must to ensure adequate time for rest and travel. Minimizing total travel cost and idle time appears to be desirable. Host nations play their games on home turf to the extent possible. Day games (9am start) are also possible, if the morning fog/mist/dew factors are not overwhelming. On the other hand, locations with consistent dew problems at night are better suited for day matches to ensure a more fair contest. Reserve days are required for games that are washed-out due to inclement weather during the knock-out stage of the tournament, and any feasible schedule must take that into account. Of course, the OR-driven Duckworth-Lewis rules for rain-interrupted games are used to maximize the chances of getting a fair contest under the circumstances.
The Big Picture
The global cricket season is itself pretty packed and the overall schedule has come in for much criticism due to player burn out as well as overselling the game. As a cricket fan, it is pretty obvious that the status-quo is so dismal that better scheduling can maximize long-term revenue while also minimizing burn-out. Overselling is a major issue, not only because it kills the golden goose (long-term fan involvement in the game), but the number of inconsequential games being played has lead to match-fixing and spot-fixing (similar to point-shaving in basketball). The money involved in illegal betting is mind-blowing, and many suspect that it already is or could become another source of income for terrorist groups operating in the Af-Pak area. One does not expect OR to help resolve all these issues, but it can certainly be applied to some of the key ones it is well suited for, and the cascading positive effects can make a difference to the overall situation.
Potential OR Approaches
The Traveling Tournament Problem (TTP) popularized by Dr. Mike Trick at CMU appears to be a good starting point for improving the schedule for future cricket tournaments, including the world cup and the IPL, and ultimately, the global cricket season itself. Exact or heuristic approaches that combine constraint programming with MIP models appear to be well-suited to such problems given the complex and 'idiosyncratic nature' of some the scheduling constraints that tend to be imposed in cricket, as well the value added by even small scheduling improvements.
Looking Forward: The 2015 World Cup and Beyond
The next edition with be the "Anzac" world cup, hosted jointly by Australia and New Zealand. The distances between some cities that are likely to host some of the matches can be enormous (e.g. Perth and Sydney are more than 4000 miles apart) or relatively tiny (e.g. intra-NZ games, or Australian east-coast games (Melbourne-Adelaide is less than 800 miles), and the value that can be obtained by adopting optimized schedules can be significant.
Conclusion
I hope this post presents a reasonable high-level overview of cricket tournament schedules and motivates interested OR'ers to further investigate this problem. As a cricket tragic and OR professional, I would be happy to contribute toward any such effort.
(To be submitted as an entry to the March-2011 Informs Blog challenge: OR and Sports)
Thursday, March 3, 2011
Low Probability High Consequence Sporting Wins
LPHC events are always interesting since the implications (conditional expected cost), given the occurrence of the event, tend to have cascading side-effects. In risk-management analytics (e.g. optimal Hazmat routing), limiting the conditional expected risk turns out to be important. On the other hand, in the realm of sport, LPHC events are often quite desirable. Who doesn't love to hear and re-hear those stories about back-to-the-wall fight-backs and come-from-nowhere wins? However, among such magic moments, only a small subset have long-term and wide-ranging implications. Often, we have to wait for years or decades to see how the story unravels and how far the ripple effects go.
Yesterday's massive sporting upset of England by Ireland in the cricket world cup - yet another (truly) sensational match in my hometown within a week - has the potential to fall under this category for a variety of reasons. First some highlights of Kevin O'Brien's amazing counter-attack.
USA beat England in the 1950 soccer world cup - an incredible upset, but one that ultimately did not make much of a difference to the sport (arguably). On the other hand, Joe Namath's 'guaranteed' win in Super-bowl III could be termed a LPHC event since it seems to have played a big part in strengthening the NFL brand by creating a gripping storyline for the future. Or more recently, the 'Invictus' rugby story of South African Springboks that united a nation on the verge of being torn apart by the after-effects of the inhuman apartheid - yes, it is a true story. Sports, like politics is local and everybody has their favorite LPHC picks. For brevity, the top two on my short-list of LPHC sporting upsets are:
1. The 1980 miracle on ice. The uplifting impact of this story on people who haven't even stepped on ice, and the inspiring convergence of the sequence of events, imho, transcends sport, nationality, and time.
2. Was the 1980s the last and greatest decade of pure sporting action around the world, across all sport? India's cricket win over the invincible Caribbean world champions in the 1983 world cup final. If I recall, the odds of team-India winning the world-cup was something like 50, 000 to 1. A eleven-person nobody-group of not particularly athletic cricketers, speaking different languages, practicing many different religions, and coming from different backgrounds, yet each proudly representing a poverty-ridden, but democratic nation of 800 million hopefuls. Competing against the supremely powerful, atrociously talented, all-conquering team that was the West Indies. Watch the trailer to the wonderful documentary film tribute to this Windies team ("Fire in Babylon").
Notice the almost subdued celebrations in those days, and compare with the over-the-top ones by Indian cricketers today!
Incredible though that day was, nobody predicted the after-effects.Events on that London day started a three-decade long perfect cricketing storm that is now a multi-billion dollar professional sports-and-media franchise with a billion-person advertising market, and growing.
Ireland as a nation is financially reeling. Many people are moving out of Ireland, bringing back memories of the potato famine days in the 19th century. Cricketing-wise, they may not even be allowed to compete in the next world cup in 2015. As if all this wasn't enough, they still have to overcome (?) the proverbial luck of the Irish. Can they do the impossible?
Yesterday's massive sporting upset of England by Ireland in the cricket world cup - yet another (truly) sensational match in my hometown within a week - has the potential to fall under this category for a variety of reasons. First some highlights of Kevin O'Brien's amazing counter-attack.
USA beat England in the 1950 soccer world cup - an incredible upset, but one that ultimately did not make much of a difference to the sport (arguably). On the other hand, Joe Namath's 'guaranteed' win in Super-bowl III could be termed a LPHC event since it seems to have played a big part in strengthening the NFL brand by creating a gripping storyline for the future. Or more recently, the 'Invictus' rugby story of South African Springboks that united a nation on the verge of being torn apart by the after-effects of the inhuman apartheid - yes, it is a true story. Sports, like politics is local and everybody has their favorite LPHC picks. For brevity, the top two on my short-list of LPHC sporting upsets are:
1. The 1980 miracle on ice. The uplifting impact of this story on people who haven't even stepped on ice, and the inspiring convergence of the sequence of events, imho, transcends sport, nationality, and time.
2. Was the 1980s the last and greatest decade of pure sporting action around the world, across all sport? India's cricket win over the invincible Caribbean world champions in the 1983 world cup final. If I recall, the odds of team-India winning the world-cup was something like 50, 000 to 1. A eleven-person nobody-group of not particularly athletic cricketers, speaking different languages, practicing many different religions, and coming from different backgrounds, yet each proudly representing a poverty-ridden, but democratic nation of 800 million hopefuls. Competing against the supremely powerful, atrociously talented, all-conquering team that was the West Indies. Watch the trailer to the wonderful documentary film tribute to this Windies team ("Fire in Babylon").
Notice the almost subdued celebrations in those days, and compare with the over-the-top ones by Indian cricketers today!
Incredible though that day was, nobody predicted the after-effects.Events on that London day started a three-decade long perfect cricketing storm that is now a multi-billion dollar professional sports-and-media franchise with a billion-person advertising market, and growing.
Ireland as a nation is financially reeling. Many people are moving out of Ireland, bringing back memories of the potato famine days in the 19th century. Cricketing-wise, they may not even be allowed to compete in the next world cup in 2015. As if all this wasn't enough, they still have to overcome (?) the proverbial luck of the Irish. Can they do the impossible?
Sunday, February 27, 2011
Analytics and Cricket - V: Death by Forecast
Is cricket the only sport in the world where a game-changing decision, i.e., an "out" is decided by a forecast? Welcome to the "leg before wicket" rule. Here, the umpire (aka the referee in the USA) gives the batsman out if he deems that the ball would have gone on to hit the stumps if the pad had not come in the way. A good percentage of outs in cricket occur via the LBW. There is some additional 'fine print' that must be satisfied in addition to the above requirement for a batsman to be given out and too technical for this tab. Here's a clip of Pakistan's great pacer Waqar Younis getting some LBWs with his fast (90-95mph+) in swinging (curve ball) yorkers in the early 1990s.
The LBW rule is unique in that a batsman is given out based on something that did not actually occur, but would have probably occurred. Not surprisingly, LBWs are the most contentious decisions in cricket. Yesterday's sensational world-cup match between India and England in my home town that ended in a dramatic last-ball tie included an LBW incident. The 2011 cricket world cup allows the use of the "Hawk-eye" trajectory predictor to help arrive at the best decision. Similar technology is used in many other sports now such as tennis. In yesterday's instance, a batsman was deemed "not out", overruling Hawk-eye's prediction, since he was struck on his pads, which was more than 2.5 metres in front of the stumps at the time of impact, a 'magic number' threshold beyond which, the Hawk-eye forecast is deemed practically unreliable. I believe that this is just the start of the problem.
Spin bowling. How the heck is Hawk-eye going to predict the amount of spin (turn) off the pitch ?? In more complex cases, there is drift, dip, as well as turn. Watch the Aussie legend Shane Warne bowl this incredible "ball of the century" some 17 years ago. Would Hawk-eye have been able to accurately predict the path of the ball after making contact with the surface of the cricket pitch ?? I some how doubt it.
Strictly speaking, forecast-based decisions do exist in other sports as well. In basketball, we have goal tends that are probabilistic calls in the sense that the scoreboard is updated based on a forecast rather than an actual basket. Are there other popular sports where non-trivial game-time decisions involve a forecast of some kind ?
In cricket, the forecasting story just doesn't stop there. We have forecast-based rules for weather-affected matches that were devised by OR professors in England - something which we already talked about before. These are causal predictive models employed regularly in professional cricket. Yet another reason why cricket is called the 'game of glorious uncertainties", and why the game of cricket is always an applied-mathematician's delight.
Given that this is world-cup cricket time, the next post will (probably) center around cricket, and will focus on a very deterministic analytical modeling element. Go India ...
The LBW rule is unique in that a batsman is given out based on something that did not actually occur, but would have probably occurred. Not surprisingly, LBWs are the most contentious decisions in cricket. Yesterday's sensational world-cup match between India and England in my home town that ended in a dramatic last-ball tie included an LBW incident. The 2011 cricket world cup allows the use of the "Hawk-eye" trajectory predictor to help arrive at the best decision. Similar technology is used in many other sports now such as tennis. In yesterday's instance, a batsman was deemed "not out", overruling Hawk-eye's prediction, since he was struck on his pads, which was more than 2.5 metres in front of the stumps at the time of impact, a 'magic number' threshold beyond which, the Hawk-eye forecast is deemed practically unreliable. I believe that this is just the start of the problem.
Spin bowling. How the heck is Hawk-eye going to predict the amount of spin (turn) off the pitch ?? In more complex cases, there is drift, dip, as well as turn. Watch the Aussie legend Shane Warne bowl this incredible "ball of the century" some 17 years ago. Would Hawk-eye have been able to accurately predict the path of the ball after making contact with the surface of the cricket pitch ?? I some how doubt it.
Strictly speaking, forecast-based decisions do exist in other sports as well. In basketball, we have goal tends that are probabilistic calls in the sense that the scoreboard is updated based on a forecast rather than an actual basket. Are there other popular sports where non-trivial game-time decisions involve a forecast of some kind ?
In cricket, the forecasting story just doesn't stop there. We have forecast-based rules for weather-affected matches that were devised by OR professors in England - something which we already talked about before. These are causal predictive models employed regularly in professional cricket. Yet another reason why cricket is called the 'game of glorious uncertainties", and why the game of cricket is always an applied-mathematician's delight.
Given that this is world-cup cricket time, the next post will (probably) center around cricket, and will focus on a very deterministic analytical modeling element. Go India ...
Sunday, February 20, 2011
Dating and wedding logistics: an OR opportunity area?
Ignighter.com is yet another online dating portal start-up by a bunch of young New Yorkers. However, this one is a bit more interesting from the OR perspective. It specifically targets 'group dating' noting that "meeting someone one-on-one is more awkward than a junior high dance" and also acts as a 'dating logistics' enabler. They also do 'group profile' matching. One gets the feeling that there's bound to be some OR methods applicable here to improve upon this original idea.
However, the most interesting aspect of this startup was that it was not very successful in igniting the NYC scene, but within a few months, their biggest customer segment came from India, much to their surprise. Why? pulling off one-one meeting coups in India generally tend to be even more daunting. Plus the fact that Indians love novelty while also holding on to the time-tested. As we noted a couple of weeks ago, traditional weddings in India makes up a huge fraction of the weddings in world at any point in time (H-hour, the most auspicious time for the final ritual, occurs at night for North Indian weddings, and during the day in the South :). This makes an upwardly mobile and economically liberalized India an irresistible growth market for practically anything new, from stealth jets to instant Mehendi/Henna, and of course, novel online dating services. To back up this claim, we go back to our tried and tested indicator - Indian movies. Yes, we have a new B'wood movie around the dating/movie logistics theme. In fact, this one is a pretty decent and successful yarn about an pair of entrepreneurs who make it big scoring contracts for the scheduling, synchronizing, and sequencing of the various events in an Indian wedding, which are among the most elaborate, yet intricate, in the world. The complexity and the number of hard and soft constraints that have to be satisfied here is likely to be challenging. Seems like a great niche O.R. opportunity area.
However, the most interesting aspect of this startup was that it was not very successful in igniting the NYC scene, but within a few months, their biggest customer segment came from India, much to their surprise. Why? pulling off one-one meeting coups in India generally tend to be even more daunting. Plus the fact that Indians love novelty while also holding on to the time-tested. As we noted a couple of weeks ago, traditional weddings in India makes up a huge fraction of the weddings in world at any point in time (H-hour, the most auspicious time for the final ritual, occurs at night for North Indian weddings, and during the day in the South :). This makes an upwardly mobile and economically liberalized India an irresistible growth market for practically anything new, from stealth jets to instant Mehendi/Henna, and of course, novel online dating services. To back up this claim, we go back to our tried and tested indicator - Indian movies. Yes, we have a new B'wood movie around the dating/movie logistics theme. In fact, this one is a pretty decent and successful yarn about an pair of entrepreneurs who make it big scoring contracts for the scheduling, synchronizing, and sequencing of the various events in an Indian wedding, which are among the most elaborate, yet intricate, in the world. The complexity and the number of hard and soft constraints that have to be satisfied here is likely to be challenging. Seems like a great niche O.R. opportunity area.
Saturday, February 12, 2011
Driving on 'E': Mad Max and OR
While doing my commute to work yesterday on the parkways of NY:
9 AM: fuel indicator hits 'E' and still some distance from the destination. After an initial surge of panic, like any good OR person i decided to build a quick inventory model of the quantity of gas (that would 'petrol' if you are Indian) held in my car. With all those NY drivers racing at breakneck speed (like those Mohawked goons from the Aussie outback in 'Mad Max'), doing real-time inventory optimization and driving safely is not easy.
9:10AM: Consciously slow down inventory depletion rate. Stay within your auto's fuel efficient range around 45-55 mph. Resist the urge to go too slow or too fast. Turn off heating - maybe that would help too.
9:12 AM: Toyota surely must have sound safety stock calculations thrown in to calibrate the fuel meter, so 'E' is more likely to be the 'replenishment point' rather than a truly empty "back order" point. Reasoning helps reduce panic. It's a new job in a new geographical area, but the car is the same old and trusted sedan.
9:15AM: Use GPS to locate nearest replenishment point. The nearest one was just 2 miles away. Good. When I get to that spot, there's just a pile of snow. Data issues with the GPS.
9:20 AM: Do I trust the GPS and search for another gas station or do I head to the office and postpone my decision to refill on return? feeling confident that there is enough in reserve, I head straight for the office staying on the highway, where cars are more fuel efficient, and make it, and park the car in the shade.
12PM: Offline analytics to plot return trip in the evening. I find this really great site. For an input car make and model, it displays the sample mean and deviation. I was not even close to riding on the edge. The distribution of gas miles after hitting 'E' shows a mean value of about 45 miles and a standard deviation of 25. Some road warriors appear to have done a hundred or more. There was one who apparently refueled his tank with 18.064 gallons and must have been running on fumes. On the other hand, the minimum value is 2.0 miles, indicating that there was a reasonable expected cost of being stranded in sub-zero conditions on a highway looking foolish and may yet be stranded in the office.
5PM: Decision optimization time. Do I bet on the analytical model and drive home to refuel at the gas station that (certainly) existed today morning next to my residence and also sells cheaper gas? or do I head for the nearest gas station from my current location? If I head home, I have roughly a 2/3 data-driven chance of making it based on the normal distribution fit to the curve on tankonempty.com. Despite this comfortable probability of success, I realize that it's relatively easier selling OR models to others. Furthermore, what happens if I'm stuck in a return-commute traffic jam? I decide to leave a little later to avoid peak traffic. However, if I can locate a gas station closest to a point on my shortest path to home, then that is an optimal route. The treacherous GPS will (hopefully) redeem itself and find a good solution to this tiny traveling salesman problem.
View Larger Map
5:45PM: The GPS is wrong. Twice in a day! I expended about 6 miles on this wild goose chase and burnt valuable daylight as well. I am no closer to home, my tank is still on 'E', my night vision is poor, and i'm freezing. I figure my odds have dropped close to coin-toss range. The GPS has been a let down as far as finding non-fictional gas stations for this highly wooded area that is still new to me. Given the darkness, I think I'm doing the sensible thing by assuming that the conditional probability of hitting a gas station given that we choose local roads (closer to residential areas), is higher. I ask Cassius to re-route me off the parkways, keeping in mind the drop in fuel efficiency on local roads, which is about 30% for my car.
6:00 PM: Despite driving at low speeds, I find a gas station within 4 miles and shell out 61$ for the 'juice'. I find that I had almost a gallon in reserve, safe and warm enough to just about take me home if I drove along the highway at optimal speeds in the first place. If only my night vision was as good as my hindsight. In the end, I was just another data blip that was pretty close to the median on the 'E' curve. Mad Max I was not. That and the fact that I don't have a spunky dog riding with me, nor a sawn-off shotgun.
9 AM: fuel indicator hits 'E' and still some distance from the destination. After an initial surge of panic, like any good OR person i decided to build a quick inventory model of the quantity of gas (that would 'petrol' if you are Indian) held in my car. With all those NY drivers racing at breakneck speed (like those Mohawked goons from the Aussie outback in 'Mad Max'), doing real-time inventory optimization and driving safely is not easy.
9:10AM: Consciously slow down inventory depletion rate. Stay within your auto's fuel efficient range around 45-55 mph. Resist the urge to go too slow or too fast. Turn off heating - maybe that would help too.
9:12 AM: Toyota surely must have sound safety stock calculations thrown in to calibrate the fuel meter, so 'E' is more likely to be the 'replenishment point' rather than a truly empty "back order" point. Reasoning helps reduce panic. It's a new job in a new geographical area, but the car is the same old and trusted sedan.
9:15AM: Use GPS to locate nearest replenishment point. The nearest one was just 2 miles away. Good. When I get to that spot, there's just a pile of snow. Data issues with the GPS.
9:20 AM: Do I trust the GPS and search for another gas station or do I head to the office and postpone my decision to refill on return? feeling confident that there is enough in reserve, I head straight for the office staying on the highway, where cars are more fuel efficient, and make it, and park the car in the shade.
12PM: Offline analytics to plot return trip in the evening. I find this really great site. For an input car make and model, it displays the sample mean and deviation. I was not even close to riding on the edge. The distribution of gas miles after hitting 'E' shows a mean value of about 45 miles and a standard deviation of 25. Some road warriors appear to have done a hundred or more. There was one who apparently refueled his tank with 18.064 gallons and must have been running on fumes. On the other hand, the minimum value is 2.0 miles, indicating that there was a reasonable expected cost of being stranded in sub-zero conditions on a highway looking foolish and may yet be stranded in the office.
5PM: Decision optimization time. Do I bet on the analytical model and drive home to refuel at the gas station that (certainly) existed today morning next to my residence and also sells cheaper gas? or do I head for the nearest gas station from my current location? If I head home, I have roughly a 2/3 data-driven chance of making it based on the normal distribution fit to the curve on tankonempty.com. Despite this comfortable probability of success, I realize that it's relatively easier selling OR models to others. Furthermore, what happens if I'm stuck in a return-commute traffic jam? I decide to leave a little later to avoid peak traffic. However, if I can locate a gas station closest to a point on my shortest path to home, then that is an optimal route. The treacherous GPS will (hopefully) redeem itself and find a good solution to this tiny traveling salesman problem.
View Larger Map
5:45PM: The GPS is wrong. Twice in a day! I expended about 6 miles on this wild goose chase and burnt valuable daylight as well. I am no closer to home, my tank is still on 'E', my night vision is poor, and i'm freezing. I figure my odds have dropped close to coin-toss range. The GPS has been a let down as far as finding non-fictional gas stations for this highly wooded area that is still new to me. Given the darkness, I think I'm doing the sensible thing by assuming that the conditional probability of hitting a gas station given that we choose local roads (closer to residential areas), is higher. I ask Cassius to re-route me off the parkways, keeping in mind the drop in fuel efficiency on local roads, which is about 30% for my car.
6:00 PM: Despite driving at low speeds, I find a gas station within 4 miles and shell out 61$ for the 'juice'. I find that I had almost a gallon in reserve, safe and warm enough to just about take me home if I drove along the highway at optimal speeds in the first place. If only my night vision was as good as my hindsight. In the end, I was just another data blip that was pretty close to the median on the 'E' curve. Mad Max I was not. That and the fact that I don't have a spunky dog riding with me, nor a sawn-off shotgun.
Saturday, February 5, 2011
Analytics, Astrology, and V-day Objectives
What predictive analytical tools do insurance companies use to manage long-term risk? The usual ones and then this. Using customer data such as month of birth, The Allstate Insurance company grouped observations into 12 buckets. To make it more fun, they labeled these buckets under their star sign ("Raashee" in Sanskrit) as a V-day joke. Then they tracked the accident levels for each of the groups. The topper in this list of trouble makers are Virgos, characterized by AllState as "worried and shy". Of course, while this was a bit of harmless fun for many people, and AllState said this was joke that fell apart, Virgos and Leos have a good reason to worry in the current economic climate, since their rates may be relatively higher because of a "pre-existing" condition. In the end, AllState went into damage-control mode and assured customers that their star-sign was never and will never be held against them :)
Assuming these results are real, it raises some interesting and entertaining questions. Is there at least some apparent correlation between your date of birth and your future on and off the road? Does the popular observation that a significant proportion of babies are conceived during the downtime in winter and thus born around August (Leo-Virgo time), have something to do these results?
When it comes to long-term decisions such as match-making, Rashee and celestial planetary alignments matter to many Indians, regardless of economic and educational levels. The recent Bollywood flick 'What's your Rashee'? comes to mind.
Indians love weddings, and a significant fraction of the marriages 'arranged' in India (which would amount to a healthy fraction of the total weddings in the world at any moment!) are based on the compatibility of horoscopes that must be determined by an expert astrologer. A 'matching algorithm' is run to determine the compatibility of horoscopes on various attributes. The outcome is an integer and a certain lower threshold must be met for the alliance to be considered worthwhile. The bigger the score, the brighter the predicted future of the proposed marriage. An unattainable upper bound for this score among mortals is a perfect 36/36 (?) which was achieved for the divine pair of Sri Rama (an Avatar of Lord Vishnu, the preserver) and his consort, Mother Sita (the daughter of Goddess Earth), whose perfect union is the basis of one of the two great Indian epics, the Ramayana.
In today's world, horoscope-matching is a fun and educative exercise for Indian couples ready to take the plunge, while also connecting with many thousand years of uninterrupted native culture. However, the problem arises when matchmakers begin to take astrological (or analytical) predictions way too seriously, and at the expense of every other reasonable consideration such as the 'content of a person's character', as the noble Dr M. L. King Jr. said. Thus it is not surprising that even the most 'secular' of Indian politicians is a fanatical follower of astrology. As members of generation-A (the analytical generation), we would love to think that we are different but things haven't changed all that much. One only needs to look at the "analytics" employed by online dating sites. The 'Analytic Age' blog had an interesting post relating to this a while ago. When it comes to making strategically useful match-making predictions, today's analytics is not much of an improvement.
Tactical Level
On the other hand, when it comes to tactics, banking on stars to bail us out on V-days and anniversaries is a recipe for 'crash and burn'. In this rare instance, a large variance can actually be good since it is the opposite of 'routine and boring', in keeping with the 'variety seeking' behavior of shoppers observed in descriptive retail analytics. But this risk is at odds with the eventual reward so we must choose our objective with care. V-days follows a geometric probability distribution, where you have to win every year just to stay in the game. One big meltdown and you are out, regardless of the big wins you had in the past when your mojo peaked. By all accounts, the expectation of tolerance on these days is ruthlessly Markovian, so remember the gambler's ruin and plan accordingly.
Operational Level
Under the assumptions of our tactical model, the default aim on the eve of any given V-day is to minimize maximum regret so we can live to see another V-day. On the other hand, if we want to go for it on 4th down, then maximizing expected value it is. However, the operational plan must have an ability to fall-back to the default objective. This way we can contain second-order effects (collateral damage) while also constraining our primary losses.
This will be a submission toward the February Informs blog challenge on 'OR and Love'.
Assuming these results are real, it raises some interesting and entertaining questions. Is there at least some apparent correlation between your date of birth and your future on and off the road? Does the popular observation that a significant proportion of babies are conceived during the downtime in winter and thus born around August (Leo-Virgo time), have something to do these results?
When it comes to long-term decisions such as match-making, Rashee and celestial planetary alignments matter to many Indians, regardless of economic and educational levels. The recent Bollywood flick 'What's your Rashee'? comes to mind.
Indians love weddings, and a significant fraction of the marriages 'arranged' in India (which would amount to a healthy fraction of the total weddings in the world at any moment!) are based on the compatibility of horoscopes that must be determined by an expert astrologer. A 'matching algorithm' is run to determine the compatibility of horoscopes on various attributes. The outcome is an integer and a certain lower threshold must be met for the alliance to be considered worthwhile. The bigger the score, the brighter the predicted future of the proposed marriage. An unattainable upper bound for this score among mortals is a perfect 36/36 (?) which was achieved for the divine pair of Sri Rama (an Avatar of Lord Vishnu, the preserver) and his consort, Mother Sita (the daughter of Goddess Earth), whose perfect union is the basis of one of the two great Indian epics, the Ramayana.
In today's world, horoscope-matching is a fun and educative exercise for Indian couples ready to take the plunge, while also connecting with many thousand years of uninterrupted native culture. However, the problem arises when matchmakers begin to take astrological (or analytical) predictions way too seriously, and at the expense of every other reasonable consideration such as the 'content of a person's character', as the noble Dr M. L. King Jr. said. Thus it is not surprising that even the most 'secular' of Indian politicians is a fanatical follower of astrology. As members of generation-A (the analytical generation), we would love to think that we are different but things haven't changed all that much. One only needs to look at the "analytics" employed by online dating sites. The 'Analytic Age' blog had an interesting post relating to this a while ago. When it comes to making strategically useful match-making predictions, today's analytics is not much of an improvement.
Tactical Level
On the other hand, when it comes to tactics, banking on stars to bail us out on V-days and anniversaries is a recipe for 'crash and burn'. In this rare instance, a large variance can actually be good since it is the opposite of 'routine and boring', in keeping with the 'variety seeking' behavior of shoppers observed in descriptive retail analytics. But this risk is at odds with the eventual reward so we must choose our objective with care. V-days follows a geometric probability distribution, where you have to win every year just to stay in the game. One big meltdown and you are out, regardless of the big wins you had in the past when your mojo peaked. By all accounts, the expectation of tolerance on these days is ruthlessly Markovian, so remember the gambler's ruin and plan accordingly.
Operational Level
Under the assumptions of our tactical model, the default aim on the eve of any given V-day is to minimize maximum regret so we can live to see another V-day. On the other hand, if we want to go for it on 4th down, then maximizing expected value it is. However, the operational plan must have an ability to fall-back to the default objective. This way we can contain second-order effects (collateral damage) while also constraining our primary losses.
This will be a submission toward the February Informs blog challenge on 'OR and Love'.
Saturday, January 29, 2011
The honest politician and other rare events
This post is mildly motivated by the INFORMS blog challenge this month dealing with 'OR and politics'. This tab dabbles with OR-eyed viewpoints of Indian political events from time to time. Past posts on OR and politics can be found here, here and here. The idea for this post arose from real issues in OR practice and business analytics. Yet, there is an interesting element of politics as well.
Consider this hypothetical experiment. We select a list of many thousand past political leaders from around the world and generate ratings on multiple attributes that provide valuable insight into their their 'level of honesty' derived from fact-driven records during their leadership tenure. On the right hand side, we have a yes/no binary indicator on whether that politician was generally considered "honest" or "dishonest". Our objective is simple: generate the probability that an input politician is honest, given a set of scores for each of his/her performance attributes.
We use a binary logit model (i.e. logistic regression) to do this and use historical data to calibrate the parameters using the maximum likelihood estimate approach. Since we have a fairly large sample size, we get a good model fit and hit all the right notes as far as confidence intervals, etc. The statistical model shows a good fit. But how well will it predict in real life? These are two different stories.
Politicians strongly rated as honest and statesmanlike are a rare species. Indian legend regards King Harishchandra as an exemplar for honesty in public life, which is not surprising given that he never uttered a single lie in his life, and greatly influenced the the first person in the next list. More recently, 'Mahatma' Gandhi, Abe Lincoln, and Nelson Mandela. In current times and keeping with contemporary mores, a Barack Obama (perhaps), Dr. Abdul Kalaam of India, or a Helen Clark of New Zealand, ..., the list of people keeping it on the level is quite short. It is likely that we will find our predictive analytical model is (far too) good as far as picking crooked politicians. If 99% of politicians are dishonest, then it is very easy to get a good fit. In fact, a 1-line model that simply returns "crooked politician" is a good one - it is 99% accurate. However, this model is not very interesting. Our focus and curiosity is driven by finding those that fall in that elusive 1%. A "NO" model fails 100% in this regard. How well did our statistically calibrated predictive model fit the "YES" instances? Most likely it did a pretty poor job and far below the expected rate of good guys. In fact, if you were very careless, your computer program may even treat some of these 'YES' data points as nuisance value/outliers! This situation is kinda like the inverse of the analytical problem of fraud detection (pun unintended). Consequently, if we fed the model, say, 'Honest' Abe Lincoln's attributes, we would be disappointed with the output. Our model moves into the domain of truthiness. On the other hand, a 'monkey model' that randomly generates answers with a mean "YES" rate of 1% may be more useful. Our challenge is to be able to do better than the monkey.
To do that, we turn to analytical work done in political science. Folks here (and in areas like new drug discovery) often work with predictive math models for rare events and some literature search in these areas indicate that there are quick (but not obvious) fixes to such plain-vanilla predictive models that we tend to use mechanically in OR projects. In particular, these corrections ensure that the natural imbalance inherent in the training data is accounted for in the right way and by the right amount.
The lesson, if any, from this experiment is that the basic act of testing predictive models on hold-out or hidden samples must never be bypassed. Fitting well to historical data is necessary for our validation, but certainly not sufficient for a customer's satisfaction. It does NOT imply "useful predictor". Not even if we have a lot of data. Furthermore, when we build a prescriptive analytical layer by embedding our predictive model within an optimization framework to determine the optimal attributes that maximize some objective, the external effects of a bad predictive model become pronounced. Optimization magnifies the silliness of a bad prediction. It literally takes it to an extreme point. In fact, an advantage of having a prescriptive layer is that it can often tell if the underlying predictive layer is playing politics with you.
Consider this hypothetical experiment. We select a list of many thousand past political leaders from around the world and generate ratings on multiple attributes that provide valuable insight into their their 'level of honesty' derived from fact-driven records during their leadership tenure. On the right hand side, we have a yes/no binary indicator on whether that politician was generally considered "honest" or "dishonest". Our objective is simple: generate the probability that an input politician is honest, given a set of scores for each of his/her performance attributes.
We use a binary logit model (i.e. logistic regression) to do this and use historical data to calibrate the parameters using the maximum likelihood estimate approach. Since we have a fairly large sample size, we get a good model fit and hit all the right notes as far as confidence intervals, etc. The statistical model shows a good fit. But how well will it predict in real life? These are two different stories.
Politicians strongly rated as honest and statesmanlike are a rare species. Indian legend regards King Harishchandra as an exemplar for honesty in public life, which is not surprising given that he never uttered a single lie in his life, and greatly influenced the the first person in the next list. More recently, 'Mahatma' Gandhi, Abe Lincoln, and Nelson Mandela. In current times and keeping with contemporary mores, a Barack Obama (perhaps), Dr. Abdul Kalaam of India, or a Helen Clark of New Zealand, ..., the list of people keeping it on the level is quite short. It is likely that we will find our predictive analytical model is (far too) good as far as picking crooked politicians. If 99% of politicians are dishonest, then it is very easy to get a good fit. In fact, a 1-line model that simply returns "crooked politician" is a good one - it is 99% accurate. However, this model is not very interesting. Our focus and curiosity is driven by finding those that fall in that elusive 1%. A "NO" model fails 100% in this regard. How well did our statistically calibrated predictive model fit the "YES" instances? Most likely it did a pretty poor job and far below the expected rate of good guys. In fact, if you were very careless, your computer program may even treat some of these 'YES' data points as nuisance value/outliers! This situation is kinda like the inverse of the analytical problem of fraud detection (pun unintended). Consequently, if we fed the model, say, 'Honest' Abe Lincoln's attributes, we would be disappointed with the output. Our model moves into the domain of truthiness. On the other hand, a 'monkey model' that randomly generates answers with a mean "YES" rate of 1% may be more useful. Our challenge is to be able to do better than the monkey.
To do that, we turn to analytical work done in political science. Folks here (and in areas like new drug discovery) often work with predictive math models for rare events and some literature search in these areas indicate that there are quick (but not obvious) fixes to such plain-vanilla predictive models that we tend to use mechanically in OR projects. In particular, these corrections ensure that the natural imbalance inherent in the training data is accounted for in the right way and by the right amount.
The lesson, if any, from this experiment is that the basic act of testing predictive models on hold-out or hidden samples must never be bypassed. Fitting well to historical data is necessary for our validation, but certainly not sufficient for a customer's satisfaction. It does NOT imply "useful predictor". Not even if we have a lot of data. Furthermore, when we build a prescriptive analytical layer by embedding our predictive model within an optimization framework to determine the optimal attributes that maximize some objective, the external effects of a bad predictive model become pronounced. Optimization magnifies the silliness of a bad prediction. It literally takes it to an extreme point. In fact, an advantage of having a prescriptive layer is that it can often tell if the underlying predictive layer is playing politics with you.
Friday, January 7, 2011
About a microanalytical startup with 'OR Inside'
Continuing with the new year theme, we take a first look at CQuotient, an analytics-driven start-up in the retail industry based in the Boston area. I just finished reading this interview on a business info site. It's interesting to hear what the founder and CEO Dr. Ramakrishnan has to say (he's also got a blog listed on the roll at the bottom right of this tab. I don't expect frequent updates for a while :). A couple of things were eye-openers.
They seem to be among the very first to sharply focus on individual customer behavior. Are we seeing some of the first practically viable applications of microanalytical (copyright, 2011 :) techniques this year ? Retail science normally thrives on aggregating individual customers into sufficiently big bunches so that the law of large numbers kicks in. Then you can reliably analyze statistical and econometric models to realistically predict and optimize based on these high-level purchase patterns. A retail microanalytical approach that drills down to the individual customer level looks pretty challenging to pull off in reality, but looking at the team assembled at CQuotient and the computing power available today, I wouldn't be surprised if they are onto something here.
Next, CQ will provide an 'optimal prescriptive' answer to a retailer. This convinces me that they have an application with "OR Inside" and their 'coolness quotient' just went up :). Rather than just dump a bunch of charts and qualitative insights on a tired, caffeine-deprived store manager-type and wish good luck, CQ seems to take it a step further and provides optimal decision recommendations to the retailer. Practical decision analytics can give you a pretty powerful edge since it can potentially eliminate or minimize a lot of costly guesswork. In the retail industry, which is characterized by wafer-thin margins and brutal competition, such OR-based innovations can be a big deal.
A minor grouse is that the word "OR" doesn't show up in the interview, but the content shows that all the good stuff is likely to be hidden inside. The scope for OR in the new world remains undiminished, especially if somebody is brave enough to dip their hands in messy data and put their money where their model is!
They seem to be among the very first to sharply focus on individual customer behavior. Are we seeing some of the first practically viable applications of microanalytical (copyright, 2011 :) techniques this year ? Retail science normally thrives on aggregating individual customers into sufficiently big bunches so that the law of large numbers kicks in. Then you can reliably analyze statistical and econometric models to realistically predict and optimize based on these high-level purchase patterns. A retail microanalytical approach that drills down to the individual customer level looks pretty challenging to pull off in reality, but looking at the team assembled at CQuotient and the computing power available today, I wouldn't be surprised if they are onto something here.
Next, CQ will provide an 'optimal prescriptive' answer to a retailer. This convinces me that they have an application with "OR Inside" and their 'coolness quotient' just went up :). Rather than just dump a bunch of charts and qualitative insights on a tired, caffeine-deprived store manager-type and wish good luck, CQ seems to take it a step further and provides optimal decision recommendations to the retailer. Practical decision analytics can give you a pretty powerful edge since it can potentially eliminate or minimize a lot of costly guesswork. In the retail industry, which is characterized by wafer-thin margins and brutal competition, such OR-based innovations can be a big deal.
A minor grouse is that the word "OR" doesn't show up in the interview, but the content shows that all the good stuff is likely to be hidden inside. The scope for OR in the new world remains undiminished, especially if somebody is brave enough to dip their hands in messy data and put their money where their model is!
Sunday, January 2, 2011
Skills for new graduates to succeed at OR practice
The first post this year is for OR students who plan to put their ideas into practice.
There is a significant ongoing transformation in the landscape of OR practice. At the end of the first ten years of this century, we see that traditional industries where OR has succeeded in the past such as airlines and logistics will continue to use OR methods. However, due to the saturation and the lack of radical breakthrough ideas, the minor incremental returns for spending R&D dollars will continue to discourage management from enhancing the science behind these OR approaches, and they will further outsource such tools to 3rd party vendors. In such a support mode, there is little that differentiates you from your competition.
On the other hand, the application of OR methods to new industries is very exciting, even lucrative. This year, OR will quietly make its way into more new industries. Most of the world (including the OR community?) will not know this, since OR is likely to remain hidden within a 'business analytics' agenda.
A new graduate who wants to practice OR should possess sound 'traditional' math and OR skills as well as the ability to work with large data sets locked in databases. You should be strong enough in your fundamentals to perform proof-of-concepts without asking your boss (typically one who cares not for OR or even knows what OR is) to shell out big $$ to buy you a new CPLEX or Gurobi license or a new SAS license to analyze patterns in data.
Familiarity with open-source tools such as COIN-OR and R will help since they are free for R&D. In such new industries, the ability to work with and analyze large volumes of messy data is perhaps more important, so being at ease there will give you an edge over non-OR types since you can 'take it all the way'. Remember, OR is an applied field that is tailor-made for analytics, and that is a powerful plus point.
A PhD would be preferable unless you are OK with being tagged as an OR-programmer/data analyst. Ability to communicate technical ideas with a non-OR audience in plain English is very, very important.
Business problems do not show up with "use OR" on it. The stalwarts of our field in the 1950s-1980s came up with original approaches that best suited the practical problem at hand, and using the best computing technology available, and these breakthroughs eventually became part of OR folklore and textbooks. OR best succeeds when it is explainable and insightful, and at times, a smart 10-line answer may just do the trick.
Finally, It's worth restating the obvious. The most important component of OR practice is that you build reliable solutions for real people who spend real $$ in a tough economy.
There is a significant ongoing transformation in the landscape of OR practice. At the end of the first ten years of this century, we see that traditional industries where OR has succeeded in the past such as airlines and logistics will continue to use OR methods. However, due to the saturation and the lack of radical breakthrough ideas, the minor incremental returns for spending R&D dollars will continue to discourage management from enhancing the science behind these OR approaches, and they will further outsource such tools to 3rd party vendors. In such a support mode, there is little that differentiates you from your competition.
On the other hand, the application of OR methods to new industries is very exciting, even lucrative. This year, OR will quietly make its way into more new industries. Most of the world (including the OR community?) will not know this, since OR is likely to remain hidden within a 'business analytics' agenda.
A new graduate who wants to practice OR should possess sound 'traditional' math and OR skills as well as the ability to work with large data sets locked in databases. You should be strong enough in your fundamentals to perform proof-of-concepts without asking your boss (typically one who cares not for OR or even knows what OR is) to shell out big $$ to buy you a new CPLEX or Gurobi license or a new SAS license to analyze patterns in data.
Familiarity with open-source tools such as COIN-OR and R will help since they are free for R&D. In such new industries, the ability to work with and analyze large volumes of messy data is perhaps more important, so being at ease there will give you an edge over non-OR types since you can 'take it all the way'. Remember, OR is an applied field that is tailor-made for analytics, and that is a powerful plus point.
A PhD would be preferable unless you are OK with being tagged as an OR-programmer/data analyst. Ability to communicate technical ideas with a non-OR audience in plain English is very, very important.
Business problems do not show up with "use OR" on it. The stalwarts of our field in the 1950s-1980s came up with original approaches that best suited the practical problem at hand, and using the best computing technology available, and these breakthroughs eventually became part of OR folklore and textbooks. OR best succeeds when it is explainable and insightful, and at times, a smart 10-line answer may just do the trick.
Finally, It's worth restating the obvious. The most important component of OR practice is that you build reliable solutions for real people who spend real $$ in a tough economy.
Thursday, December 30, 2010
Analytics and Cricket - IV: The Great Indian Coin Toss
We'll end this year with the humblest of analytical models - the coin toss. It is an important benchmark. After all, if your predictive business model can consistently outperform a coin-toss approach, then that could be a big deal in many practical situations. So what do we make of the Indian cricket captain Mahendra Singh Dhoni's (MS for short) performance with the coin? He's lost 13 of the last 14 trials!
A coin toss can be a big deal in cricket, since a 'win' allows you to decide whether to bat or bowl first. A 'flat' wicket means it's a great one to bat on and make best use of it, and the opposition gets to play on the same pitch after potential wear and tear. A 'sticky wicket' or overcast conditions on a 'green' pitch means bowling first could be a great option since batting will be difficult for the first few hours due to the 'swing' and 'seam' movement potentially available to the bowlers.
Die-hard cricket fans like me and players are among the most superstitious in the world due to the long and complex nature of the game. MS gets blamed for "losing" the toss and he's even asked for tips on improving his record :) Useful analytical models are nice to have, but they could go horribly wrong, especially when applied to cricket ... Before the sports fan begins to question his faith in science and even doubt the fundamental idea of Bernoulli trials and the law of large numbers, we note that if MS had lost 14 tosses in a row, that would have been an extreme "achievement" since the probability of that happening would have been roughly 60 in a million, and that did not happen. Phew! that counts as favorable evidence.
With the India-South Africa cricket series tied at 1-1, and with one test match to go, we have no choice but to seek solace in the scientific estimate that our fearless captain still has a 50% chance of winning the toss in Cape Town. I know that doesn't sound encouraging. But there's got to be a point in time when nature is going to bring that win-loss average back close to 50%. Will that happen in 2011? who knows ...
In 2010, MS had several ways of losing 13 of the 14 tosses. More simply, he had 14 ways of winning exactly one toss. We know that the probability of winning m of n tosses follows the Binomial distribution, and we can find out online here that the chance of losing 13 out of 14 is still tiny, at 0.00085. In other words, the chance of him winning 2 or more tosses in 2010 was greater than 999/1000, and yet that did not happen!
Like most great teams, this current Indian cricket team does not depend much on the outcome of the toss. Put into bat on a green, bouncy wicket under overcast conditions, they still managed to defeat RSA in 4 days and displayed amazing skill and resilience in the process. Still, it wouldn't hurt to begin the final match between No.1 in the world (India) and No.2 in the world (RSA), starting on Jan 2, 2011, by winning the coin-toss. If MS loses that toss, then the probability of this extended streak over 15 trials would be around 0.0004, i.e., 50% less than the already dismal number he is at today. Surely, that's unlikely, right? Let's see. What is the probability of the sequence that ends with him winning the toss on Jan 2, i.e., the chance that he wins exactly 1 of the first 14, and then win the 15th? Sadly, that's not very different. Delving into the past does not help the Indian sports fan, and talking to statisticians would not help since none of them wants to see such a rare streak end :)
It is better to look forward to the new year, where 2010 is done and dusted. We can say it again: MS has a 50% of winning the next toss, and relatively speaking, that looks so much more promising and simpler to comprehend.
Happy New Year and Go India!
A coin toss can be a big deal in cricket, since a 'win' allows you to decide whether to bat or bowl first. A 'flat' wicket means it's a great one to bat on and make best use of it, and the opposition gets to play on the same pitch after potential wear and tear. A 'sticky wicket' or overcast conditions on a 'green' pitch means bowling first could be a great option since batting will be difficult for the first few hours due to the 'swing' and 'seam' movement potentially available to the bowlers.
Die-hard cricket fans like me and players are among the most superstitious in the world due to the long and complex nature of the game. MS gets blamed for "losing" the toss and he's even asked for tips on improving his record :) Useful analytical models are nice to have, but they could go horribly wrong, especially when applied to cricket ... Before the sports fan begins to question his faith in science and even doubt the fundamental idea of Bernoulli trials and the law of large numbers, we note that if MS had lost 14 tosses in a row, that would have been an extreme "achievement" since the probability of that happening would have been roughly 60 in a million, and that did not happen. Phew! that counts as favorable evidence.
With the India-South Africa cricket series tied at 1-1, and with one test match to go, we have no choice but to seek solace in the scientific estimate that our fearless captain still has a 50% chance of winning the toss in Cape Town. I know that doesn't sound encouraging. But there's got to be a point in time when nature is going to bring that win-loss average back close to 50%. Will that happen in 2011? who knows ...
In 2010, MS had several ways of losing 13 of the 14 tosses. More simply, he had 14 ways of winning exactly one toss. We know that the probability of winning m of n tosses follows the Binomial distribution, and we can find out online here that the chance of losing 13 out of 14 is still tiny, at 0.00085. In other words, the chance of him winning 2 or more tosses in 2010 was greater than 999/1000, and yet that did not happen!
Like most great teams, this current Indian cricket team does not depend much on the outcome of the toss. Put into bat on a green, bouncy wicket under overcast conditions, they still managed to defeat RSA in 4 days and displayed amazing skill and resilience in the process. Still, it wouldn't hurt to begin the final match between No.1 in the world (India) and No.2 in the world (RSA), starting on Jan 2, 2011, by winning the coin-toss. If MS loses that toss, then the probability of this extended streak over 15 trials would be around 0.0004, i.e., 50% less than the already dismal number he is at today. Surely, that's unlikely, right? Let's see. What is the probability of the sequence that ends with him winning the toss on Jan 2, i.e., the chance that he wins exactly 1 of the first 14, and then win the 15th? Sadly, that's not very different. Delving into the past does not help the Indian sports fan, and talking to statisticians would not help since none of them wants to see such a rare streak end :)
It is better to look forward to the new year, where 2010 is done and dusted. We can say it again: MS has a 50% of winning the next toss, and relatively speaking, that looks so much more promising and simpler to comprehend.
Happy New Year and Go India!
Wednesday, December 8, 2010
The shortest path between OR jobs
Driving from my old job in the Burlington, MA area to Elmsford, NY (near my new job location at Yorktown Heights) took less than 3 hours. It seemed like a race-course full of caffeine-high jihadi drivers after all those leisurely strolls through the somnolent country roads of Maine. I got a newer GPS product (yet another Garmin) from an e-tailer. Having worked on retail pricing during the past four years, and this being a pre-Black Friday deal, I almost reflexively asked for a price match and sure enough - there was 60$ in savings to be had after pushing against some soft constraints. Since this was Garmin's latest version in the series it wasn't discounted on BF, so it turned out to be a pretty decent deal in the end.
The GPS product, on its short maiden voyage from MA to NY decided to take me through no fewer than four interstate highways: I-95, I-90, I-91, and I-84. The route seemed simpler on paper. I quickly realized that newer does not necessarily mean better. The re-routing algorithm is still ancient even though the newer one allegedly takes traffic congestion into account. To avoid extensive re-calculation of the shortest path in real time, the product continues to merely finds the quickest way to get back to plan. This is an approach typically used in airline online crew recovery ops (even though fancier global optimization algorithms have been available on paper). In general, this is not a bad idea as long as you don't wander off deep into the reservation. Forcing a recalculation enables you to recover the faster (optimal?) route, and my ETA dropped by about 10 minutes. The newer version has an "EcoRoute" option that allows you to find minimal cost paths, in addition to the standard metrics based on distance and time. Looks like you can also plan a trip having multiple intermediate nodes. That looks like a nice TSP structure. An analysis of these new features makes for an interesting post on another day.
The GPS product, on its short maiden voyage from MA to NY decided to take me through no fewer than four interstate highways: I-95, I-90, I-91, and I-84. The route seemed simpler on paper. I quickly realized that newer does not necessarily mean better. The re-routing algorithm is still ancient even though the newer one allegedly takes traffic congestion into account. To avoid extensive re-calculation of the shortest path in real time, the product continues to merely finds the quickest way to get back to plan. This is an approach typically used in airline online crew recovery ops (even though fancier global optimization algorithms have been available on paper). In general, this is not a bad idea as long as you don't wander off deep into the reservation. Forcing a recalculation enables you to recover the faster (optimal?) route, and my ETA dropped by about 10 minutes. The newer version has an "EcoRoute" option that allows you to find minimal cost paths, in addition to the standard metrics based on distance and time. Looks like you can also plan a trip having multiple intermediate nodes. That looks like a nice TSP structure. An analysis of these new features makes for an interesting post on another day.
Wednesday, November 24, 2010
2G scam followup: True opportunity cost of misallocating scarce resources
Please see the most recent post for the preliminary analysis of this scam. This is a follow up tab posting. Per this article on rediff.com:
" .. The Comptroller & Auditor General has calculated in his official report that the exchequer lost the truly mind-boggling sum of Rs 176,645 crore (Rs 176.64 billion) .. "
So in case there was any well-intentioned doubt that the 1.76*10^12 number was cooked up, it is now very clear that this number is (sadly) official. Actually, i would expect the number to be even higher, when you compare the true opportunity cost (due to a miserably and deliberately bad mis-allocation) relative to the value of optimal allocation.
When we read about scams like this, we realize how important it is that solid OR models be built to perform exploratory studies and simulations be run prior to allocating almost priceless resources. The supreme court of India said that "the 2G scam puts all other scams [in the history of India] to shame". When so many in India are dying of starvation and are homeless, such giga-squandering of public money by a corrupt government is nothing short of a 'monetary holocaust'.
It must be made mandatory for governments and public organizations at any level to conduct an appropriate OR analysis before allocating any scarce resource that belongs to the public. If the government of India had funded an OR group to spent a exaggerated and gigantic (or microscopic if u compare with the final loss) sum of 10 million $ for an OR analytical study, it would have paid for itself many, many times over. Well-run OR projects typically cost much less while providing incredibly impressive value measured in terms of incremental-benefit/project-cost return ratios (read the Woolsey papers for more on this).
Side note
Statistically, #barkhagate is turning out to be the most continually tweeted phrase in virtual India. Ever. It is trending so hot, you can make a virtual omelet there. Social media is making its presence felt in a very real way wrt real world issues in the largest democracy in the world, and consequently, the manipulative mainstream English media in India that had previously closed ranks on this topic so far, is now being forced to cover this critical news.
" .. The Comptroller & Auditor General has calculated in his official report that the exchequer lost the truly mind-boggling sum of Rs 176,645 crore (Rs 176.64 billion) .. "
So in case there was any well-intentioned doubt that the 1.76*10^12 number was cooked up, it is now very clear that this number is (sadly) official. Actually, i would expect the number to be even higher, when you compare the true opportunity cost (due to a miserably and deliberately bad mis-allocation) relative to the value of optimal allocation.
When we read about scams like this, we realize how important it is that solid OR models be built to perform exploratory studies and simulations be run prior to allocating almost priceless resources. The supreme court of India said that "the 2G scam puts all other scams [in the history of India] to shame". When so many in India are dying of starvation and are homeless, such giga-squandering of public money by a corrupt government is nothing short of a 'monetary holocaust'.
It must be made mandatory for governments and public organizations at any level to conduct an appropriate OR analysis before allocating any scarce resource that belongs to the public. If the government of India had funded an OR group to spent a exaggerated and gigantic (or microscopic if u compare with the final loss) sum of 10 million $ for an OR analytical study, it would have paid for itself many, many times over. Well-run OR projects typically cost much less while providing incredibly impressive value measured in terms of incremental-benefit/project-cost return ratios (read the Woolsey papers for more on this).
Side note
Statistically, #barkhagate is turning out to be the most continually tweeted phrase in virtual India. Ever. It is trending so hot, you can make a virtual omelet there. Social media is making its presence felt in a very real way wrt real world issues in the largest democracy in the world, and consequently, the manipulative mainstream English media in India that had previously closed ranks on this topic so far, is now being forced to cover this critical news.
Subscribe to:
Posts (Atom)