Thursday, March 23, 2023
Temporary note
I am posting this to demonstrate that this account is associated with btilly@gmail.com.
I'll delete it soon, so don't get attached to it.
Thursday, January 7, 2016
Why NSA surveillance scares me
Nothing To Hide explains why you should care about surveillance even though you have nothing to hide. But it missed the point that most concerns me.
Surveillance states target people who are disliked by the established political order. Starting with politicians who might inconvenience those in charge of surveillance. This distorts the political system in a way that should scare everyone.
Here are two examples from fairly recent US history:
In my lifetime, the total number of Americans killed or injured on US soil by terrorism and related causes is under 5000. (Here is a list of incidents.) There are over 300 million Americans alive and my life is more than half-over, so my future risk of injury or death from a terrorist attack should be under 1/60,000.
Let's analyze the other side conservatively. No government in history has ever survived 3000 years. I doubt that the US government will be an exception. Surveillance sliding into dictatorship is one of the most common ways that democracies fall. So it is reasonable to give us greater than a 1% chance of falling into dictatorship in any random period of 30 years. Such as the rest of my life. I have a habit of speaking my mind in public, just as I'm doing here. Suppose this gives me at least a 10% chance of showing up on some surveillance list. Let's further suppose that being on that list gives me at least a 10% chance of having bad things happen to me once the government falls. I now have greater than a 1/10,000 chance of death or injury in my lifetime from our surveillance state.
Therefore our surveillance state is many times more likely to result in my death or injury than the terrorists which it supposedly protects me from. Both possibilities are unlikely, but surveillance is the much bigger risk.
Stop and think about that. I don't do much wrong. I'm not out to overthrow or hurt anyone. I may not like our politicians, but I bear them little ill will. I believe that the vast majority of people who are part of our security structure have the best of intentions. I believe that most of our NSA and military are truly devoted to protecting people like me. But I still conclude that I am many times more likely to suffer harm from NSA surveillance than I am from Muslim terrorists.
This fact scares me. Maybe you should be scared too.
Surveillance states target people who are disliked by the established political order. Starting with politicians who might inconvenience those in charge of surveillance. This distorts the political system in a way that should scare everyone.
Here are two examples from fairly recent US history:
- J. Edgar Hoover maintained control of the FBI and its predecessor from 1924 to 1974 because politicians knew he had the goods on them. So nobody wanted to challenge his authority, no matter what they thought of him.
For example, he ignored the Mafia until The Apalachin Meeting was discovered by local police and reported nationally by newspapers in 1957. Even then, organized crime was prioritized below his private anti-Civil Rights Movement vendetta. Even though this put his policies at odds with several Presidents.
What new Hoover will emerge in the future, and what political agenda will he impose on all of us?
- Richard Nixon is one of the most influential Presidents in US history. His accomplishments range from ending the Vietnam war to opening relations with China to creating the EPA. But he is remembered for Watergate. Why?
The Watergate scandal was about Nixon conducting surveillance against political opponents, then covering up evidence after it was discovered. The Washington Post traced a minor burglary back to Nixon, he was indicted by Congress, and then resigned on August 9 1974 once was clear that The Senate would uphold the indictment. He was re-elected in one of the largest landslides in US history not long before. He was widely hated for a generation after.
Public opinion and our politicians responded to his case with fear that his actions put us on the path to dictatorship. Memories of Germany and WW II were still fresh, and nobody wanted that to happen here.
Nixon wouldn't have been caught today. Current NSA infrastructure gives many thousands of analysts access everything that Nixon wanted, with no clues from which they can be detected outside the system. And there is no hope that they would be caught from inside the system, either. Nixon got the active cooperation of the FBI. But let's assume a rogue President didn't manage that, he only subverted one rogue analyst. According to NSA documents, there is a "we trust you'll stop" system to keep analysts from spying in the US. Analysts don't even file a report if they have done it. According to other revelations, pre-Snowden the NSA had no other effective oversight mechanism. The most likely result of an effective oversight mechanism is the possibility of embarrassment for the NSA, so they probably haven't created one since.
Will our political system get subverted? Has it been already? How would we know?
In my lifetime, the total number of Americans killed or injured on US soil by terrorism and related causes is under 5000. (Here is a list of incidents.) There are over 300 million Americans alive and my life is more than half-over, so my future risk of injury or death from a terrorist attack should be under 1/60,000.
Let's analyze the other side conservatively. No government in history has ever survived 3000 years. I doubt that the US government will be an exception. Surveillance sliding into dictatorship is one of the most common ways that democracies fall. So it is reasonable to give us greater than a 1% chance of falling into dictatorship in any random period of 30 years. Such as the rest of my life. I have a habit of speaking my mind in public, just as I'm doing here. Suppose this gives me at least a 10% chance of showing up on some surveillance list. Let's further suppose that being on that list gives me at least a 10% chance of having bad things happen to me once the government falls. I now have greater than a 1/10,000 chance of death or injury in my lifetime from our surveillance state.
Therefore our surveillance state is many times more likely to result in my death or injury than the terrorists which it supposedly protects me from. Both possibilities are unlikely, but surveillance is the much bigger risk.
Stop and think about that. I don't do much wrong. I'm not out to overthrow or hurt anyone. I may not like our politicians, but I bear them little ill will. I believe that the vast majority of people who are part of our security structure have the best of intentions. I believe that most of our NSA and military are truly devoted to protecting people like me. But I still conclude that I am many times more likely to suffer harm from NSA surveillance than I am from Muslim terrorists.
This fact scares me. Maybe you should be scared too.
Tuesday, August 13, 2013
Becoming a web designer
I recently was having a conversation with an art student who wanted a better future than her job as a barista was likely to provide. I suggested a job as a web designer. She had no idea what would be involved, so I promised to talk to a web designer that I knew then get back to her with advice about what it would take.
He laid out a very concrete guide to follow for someone starting with an art background, talent for composition, proficiency with photoshop, and a willingness to learn. I'm putting it up as a blog in case other people have suggestions, corrections, or can benefit from the advice. Here is the guide.
- Design your own website in Photoshop. What the website is does not matter, but making it look good does. Think up something obvious (for instance a gallery for your artwork), figure out what pages it will have, what they should look like, then make prototypes in Photoshop. If you've got artistic skills and know Photoshop well, this hopefully is fairly easy.
- Build that website in WordPress or another CMS (Content Management System). Now you know what you want the web page to look like, now try to make a web page that looks at it. You'll need to buy hosting, buy a domain, then install your CMS. Now try to make WordPress serve up pages that look like the ones that you have designed. You will encounter a lot of problems. Google searches will often help you find solutions. Browse http://wp.smashingmagazine.com/ for more ideas. And if you need to compromise on your initial design because you can't figure out how to make it work, do so. The goal here is to get something done that looks nice, and not necessarily to perfectly realize your first vision.
- Start freelancing. Once you have demonstrated that you can build an attractive website, you have a skill that people are willing to pay for. Go to https://www.elance.com/ and start offering your skills as a web designer. You should start at the low end of the free-lance market, say around $50/hour. You won't get steady work, but it is a nice side income. And, more importantly, you're building your experience and a portfolio.
- Look for full time work as a web designer. In a few months you should have 6-8 actual websites under your belt, and will have learned a lot more about working with customers. This is a real portfolio to point to during a search for full-time work as a web designer. Depending on luck, opportunities, location, etc, you might land the job you want, or may need to take an intern to full-time route. A reasonable salary in the Los Angeles area to aim for with the full time job is probably around $70K to do design and HTML. That varies widely over time, by geographical market, by company, etc.
- Upgrade your skills. Once you have a job, you've turned skill at composition, knowledge of photoshop, and a willingness to learn into a real career. But you have just begun. There is a lot more to learn. In particular over the next few years you need to improve on the underlying technologies of HTML, CSS, and JS. You should be trying to educate yourself about UX. You need to build a professional network. There is a long ladder to climb, but at least you've got your foot on the bottom rung.
This outline was recommended by one person that I know who became a web designer. It sounded reasonable to me, but I am not a web designer. Starving art students may be surprised at how quickly you go from designing something for yourself to getting paid significant money for your time. That happens because you're solving real problems. Your combination of artistic talents and willingness to tackle new things is worth far more to employers than either is alone.
Please share if you find this helpful, have alternate suggestions, experience, etc.
Wednesday, November 21, 2012
Speculating about the Hyperloop
Elon Musk has been dropping hints about his Hyperloop idea. (We cannot call it a proposal because he has not actually proposed it yet.) There is a lot of curiosity about what it might be. Given Elon's history, the idea will sound audacious and yet will actually be workable.
Jacques Mattheij recently speculated on the topic. His proposal has the serious problem in that the friction of the air on the sides of the tunnel would lose way too much energy. But it got me thinking, and I have what may be a more realistic proposal.
First, imagine a tube that goes in a loop from Los Angeles to San Francisco. Let's put flaps on the walls. When they are open, air pressure can equalize. When they are closed, they don't leak much. Now let's put large, heavy objects going around and around the loop. For lack of a better name, let's call them plungers. The plungers can be floated and moved very efficiently with maglev technology. As each plunger approaches, the flaps open so that air can get pushed out, then closes so that it doesn't come back in. This is not an evacuated tube (Elon explicitly says that his technology isn't an evacuated tube), but results in a decent vacuum away from the shockwaves in front of each of plunger. That eliminates most of your friction losses. I don't know how low, but Elon claims low enough that solar panels on top of the device provide more than enough energy to keep it permanently going. I see no reason to disbelieve that solar panels could do that.
Now where do the people fit in? People go into vehicles that I'll call cars, even though they aren't really cars. These cars can be fired by a railgun to match speeds with the tube, and injected in front of a plunger. We can build the plunger with a space in its front that the car fits in. This space has air trapped in it by the shock wave and so the people can breathe. On its own that space would heat up due to the friction on the gas, but you can put a heat sink (eg a block of ice) in the car and keep it comfortable inside. Near the end of your journey the plunger ejects the car from this space on a course that launches it out of the tube while the plunger continues on its way. The car is then stopped with regenerative braking that recovers most of the launch energy, resulting in surprisingly little energy loss for taking the trip.
Now what would some of the specs be? Well, Elon claims 30 minutes from downtown Los Angeles to downtown San Francisco. According to google maps that's 382 miles, which is about 600 km. So the loop should be going around 1200 km/hour. If we put 12 plungers on the loop, and have plenty of vehicles, then you get in your vehicle and have a launch opportunity every 5 minutes. Increase the number of plungers, and the time to launch can be decreased while the capacity of the system can be increased. If we space the plungers a third of a km apart, we would have 3600 of them and could be launching every second into the system. It probably is more efficient if you instead make plungers larger so that cars carry more people. So instead of car think "bus". But after the initial system is built, you can later add new entrance/exit ramps and ramp up capacity. As Elon has promised, you would not need to reserve tickets - you'd pretty much arrive and then go.
Elon also claims that his system could store a lot of energy, enough to collect energy during the day and run off of it at night. The obvious place to store energy is in the kinetic energy of the plungers. How much energy are we talking? Well 1200 km/hour is a third of a km per second. That's 55.5 kJ of energy per kg of plunger. So 64,800 kg (about 143,000 pounds) of plunger is a megawatt-hour of power. Suppose that is one plunger. If you've got a thousand plungers, and each is storing a full megawatt-hour, you could permanently consume 50 megawatts of energy. You'd use up over half of it at night, then regain it during the day. Trips taken in the early morning commute might take 50% longer than during the evening, but it is doable. This gets better if we make plungers bigger, have a more efficient system, or have more plungers. I'm sure that Elon has thought about the ideal parameters. But if heavy plungers are good, well, put in enough metal for maglev to work and then add rock. You'd store a lot of energy.
Heck, the solar panel angle is fun but not really necessary. From the point of view of the electric power grid, it would be very, very good to have a large energy sink that can even out power fluctuations. Renewable energy sources often arrive at different times than we'd like to get power out. Sometimes, like with wind, we get very sharp spikes that we need to even out. If designed properly, the Hyperloop can absorb pretty much any power spike, and can bleed enough power out to be interesting. Therefore if the power utilities are smart then they should be willing to pay to add more plungers. Not because they care about the fact that they are improving peak capacity and reducing waiting time, but because they want to be able to store more power in it.
I'm sure that there are many improvements on this design that I have not thought of but which Elon has. I'm also sure that Elon has detailed blueprints that take this from a half-baked concept to something you can start to put cost estimates on. But this idea looks doable to me, and looks like it could - at least in principle - justify all of the claims that Elon has been making for the Hyperloop.
Finally I'd love to see this built. I'd love to see it built in California. But, unless someone like Elon pushes it, I'd be willing to bet that the Chinese get it first.
BTW for further discussion, see Hacker News
Jacques Mattheij recently speculated on the topic. His proposal has the serious problem in that the friction of the air on the sides of the tunnel would lose way too much energy. But it got me thinking, and I have what may be a more realistic proposal.
First, imagine a tube that goes in a loop from Los Angeles to San Francisco. Let's put flaps on the walls. When they are open, air pressure can equalize. When they are closed, they don't leak much. Now let's put large, heavy objects going around and around the loop. For lack of a better name, let's call them plungers. The plungers can be floated and moved very efficiently with maglev technology. As each plunger approaches, the flaps open so that air can get pushed out, then closes so that it doesn't come back in. This is not an evacuated tube (Elon explicitly says that his technology isn't an evacuated tube), but results in a decent vacuum away from the shockwaves in front of each of plunger. That eliminates most of your friction losses. I don't know how low, but Elon claims low enough that solar panels on top of the device provide more than enough energy to keep it permanently going. I see no reason to disbelieve that solar panels could do that.
Now where do the people fit in? People go into vehicles that I'll call cars, even though they aren't really cars. These cars can be fired by a railgun to match speeds with the tube, and injected in front of a plunger. We can build the plunger with a space in its front that the car fits in. This space has air trapped in it by the shock wave and so the people can breathe. On its own that space would heat up due to the friction on the gas, but you can put a heat sink (eg a block of ice) in the car and keep it comfortable inside. Near the end of your journey the plunger ejects the car from this space on a course that launches it out of the tube while the plunger continues on its way. The car is then stopped with regenerative braking that recovers most of the launch energy, resulting in surprisingly little energy loss for taking the trip.
Now what would some of the specs be? Well, Elon claims 30 minutes from downtown Los Angeles to downtown San Francisco. According to google maps that's 382 miles, which is about 600 km. So the loop should be going around 1200 km/hour. If we put 12 plungers on the loop, and have plenty of vehicles, then you get in your vehicle and have a launch opportunity every 5 minutes. Increase the number of plungers, and the time to launch can be decreased while the capacity of the system can be increased. If we space the plungers a third of a km apart, we would have 3600 of them and could be launching every second into the system. It probably is more efficient if you instead make plungers larger so that cars carry more people. So instead of car think "bus". But after the initial system is built, you can later add new entrance/exit ramps and ramp up capacity. As Elon has promised, you would not need to reserve tickets - you'd pretty much arrive and then go.
Elon also claims that his system could store a lot of energy, enough to collect energy during the day and run off of it at night. The obvious place to store energy is in the kinetic energy of the plungers. How much energy are we talking? Well 1200 km/hour is a third of a km per second. That's 55.5 kJ of energy per kg of plunger. So 64,800 kg (about 143,000 pounds) of plunger is a megawatt-hour of power. Suppose that is one plunger. If you've got a thousand plungers, and each is storing a full megawatt-hour, you could permanently consume 50 megawatts of energy. You'd use up over half of it at night, then regain it during the day. Trips taken in the early morning commute might take 50% longer than during the evening, but it is doable. This gets better if we make plungers bigger, have a more efficient system, or have more plungers. I'm sure that Elon has thought about the ideal parameters. But if heavy plungers are good, well, put in enough metal for maglev to work and then add rock. You'd store a lot of energy.
Heck, the solar panel angle is fun but not really necessary. From the point of view of the electric power grid, it would be very, very good to have a large energy sink that can even out power fluctuations. Renewable energy sources often arrive at different times than we'd like to get power out. Sometimes, like with wind, we get very sharp spikes that we need to even out. If designed properly, the Hyperloop can absorb pretty much any power spike, and can bleed enough power out to be interesting. Therefore if the power utilities are smart then they should be willing to pay to add more plungers. Not because they care about the fact that they are improving peak capacity and reducing waiting time, but because they want to be able to store more power in it.
I'm sure that there are many improvements on this design that I have not thought of but which Elon has. I'm also sure that Elon has detailed blueprints that take this from a half-baked concept to something you can start to put cost estimates on. But this idea looks doable to me, and looks like it could - at least in principle - justify all of the claims that Elon has been making for the Hyperloop.
Finally I'd love to see this built. I'd love to see it built in California. But, unless someone like Elon pushes it, I'd be willing to bet that the Chinese get it first.
BTW for further discussion, see Hacker News
Labels:
Elon Musk,
Hyperloop,
power grid,
speculation,
transportation
Monday, October 29, 2012
A/B testing scale cheat sheet
This is not a guide to how to do A/B testing. If you want that, see Effective A/B Testing, or any number of companies that will help you with A/B testing. Instead this is a cheat sheet of basic facts on A/B testing (mostly on the scale involved) to help people who are beginning figure out what is feasible.
- If you've never tested, expect to find a number of 5-20% wins.
In my experience, most companies find several changes that each add in the neighborhood of 5-20% to the bottom line in their first year of testing. If you have enough volume to reliably detect wins of this size in a reasonable time frame, you should start now. If not, then you're not a great candidate for it..yet.
- Experienced testers find smaller wins.
When you first start testing, you pick up the low-lying fruit. The rapid increase in profits can spoil you. Over time the wins get smaller. And more specific to your business. Which means they will take longer to detect. Expect this. But if you're still finding an average of one 2% win every month, that is around a 25% improvement in conversion rates per year.
- What works for others likely works for you. But not always.
Companies that test a lot which have settled on simple value propositions, streamline their signup process, put down big calls to action, places the same call to action in multiple places, and do email marketing. Those are probably going to be good things for you to do as well. However Amazon famously relies on customer ratings and reviews. If you do not have their scale or product mix, you'll likely get very few reviews per product, and may get overwhelmingly negative reviews. So borrow, but test.
- Your testing methodology needs to be appropriate to your scale and experience.
A company with 5k prospects per day might like to run a complex multivariate test all of the way to signed up prospect, and be able find subtle 1% conversion wins. But they don't generate nearly enough data to do this. They may need to be satisfied with simple tests on the top step of the conversion funnel. But it would be a serious mistake for a company like Amazon to settle for such a poor testing methodology. In general you should use the most sophisticated testing methodology that you generate enough data to carry off.
- Back of the envelope: A 10% win typically needs around 800-1500 successes per version to be seen.
One of the top questions people have is how long a test takes to run. Unfortunately it depends on all sorts of things, including your traffic volume, conversion rates, the size of the win to be found, luck, and what confidence you cut the test off at. But if one version gets to 800 successes, when the new one is at 880, you can convert at a 95% confidence level. If you wait until you have 1500 versus 1650, you can convert at a 99% confidence level. This data point, combined with your knowledge of your business, gives you a starting point for planning.
- Back of the envelope: Sensitivity scales as the square root of the number of observations.
For example a 5% win takes about 2x as much sensitivity as a 10% win, which means 4x as much data. So you need 3200-6000 successes per version to see it.
- Data required is roughly linear with number of versions.
Running more versions requires a bit more data per version to reach confidence. But not a lot. Thus the amount of data you need is roughly proportional to the number of versions. (But if some versions are real dogs, it is OK to randomly move people from those versions to other versions, which speeds up tests with a lot of versions.) Before considering a complicated multivariate test, you should do a back of the envelope to see if it is feasible for your business.
- Even if you knew the theoretical win, you can't predict how long it will actually take to within a factor of 3.
An A/B test reaches confidence when the observed difference is bigger than chance alone can plausibly explain. However your observed difference is the underlying signal plus a chance component. If the chance component is in the same direction as the underlying signal, the test finishes very fast. If the chance component is the opposite direction, then you need enough data that the underlying signal overrides the chance signal, and goes on to still be larger than chance could explain. The difference in time is usually within a factor of 3 either way, but it is entirely luck which direction you get. (The rough estimates above are not too far from where you've got a 50% chance of having an answer.)
- The lead when you hit confidence is not your real win.
This is the flip side of the above point. It matters because someone usually has the thankless task of forecasting growth. If you saw what looked like an 8% win, the real win could easily be 4%. Or 12%. Nailing that number down with any semblance of precision will take a lot more data, which means time and lost sales. There generally isn't much business value in knowing how much better your better version is, but whoever draws up forecasts will find the lack of a precise answer inconvenient.
- Test early, test often.
Suppose that you have 3 changes to test. If you run 3 tests, you can learn 3 facts. If you run one test with all three changes, you don't know which change actually made a difference. Small and granular tests therefore do more to sharpen your intuition about what works.
- Testing one step in the conversion funnel is OK only if you're small and just beginning testing.
Every business has some sort of conversion funnel which potential customers go through. They see your ad, click on it, click on a registration link, actually sign up, etc. As a business, you care about people who actually made you money. Each step loses people. Generally, whatever version pushes more people through the top step gets more business in the end. Particularly if it is a big win. But not always! Therefore if testing eventual conversions takes you too long, and you're still finding 10%+ wins at the top step in your funnel, it makes business sense to test and run with those preliminary answers. You'll make some mistakes, but you'll get more right than wrong. Testing poorly is better than not testing at all.
- People respond to change.
If you change your email subject lines, people may be 2-5% more likely to click just because it is different, whether or not it is better. Conversely moving a button on the website may decrease clicks because people don't see the button where they expect it. If you've progressed to looking for small wins, then you should keep an eye out for tests where this is likely to matter, and try to dig a little deeper on this.
- A/B testing revenue takes more data. A lot more.
How much more depends on your business. But be ready to see data requirements rise by a factor of 10 or more. Why? In the majority of companies, a fairly small fraction of customers spend a lot more than average. The detailed behavior of this subgroup matters a lot to revenue, so you need enough data to average out random fluctuations in this slice of the data.
- Interaction effects are likely ignorable.
Often people have several things that they would like to test at the same time. If you have sufficient data, of course, you would like to look at each slice caused by a combination of possible versions separately, and look for interaction effects that might arise with specific combinations. In my experience, most companies don't have enough volume to do that. However if you assign people to test versions randomly, and apply common sense to avoid obvious interaction effects (eg red text on red background would cause an interaction effect), then you're probably OK. Imperfect testing is better than not testing, and the imperfection of proceeding is generally pretty small.
As always, feedback is welcome. I have helped a number of companies in the Los Angeles area on A/B testing, and this tries to the most common questions that I've encountered about how much work it is, and what returns they can hope for.
Wednesday, October 17, 2012
My son's flashcard routine
My 7 year old son is in grade 2. In the previous grade, despite his intelligence, he was significantly behind his class in handwriting, letter reversals, and spelling. He was getting extra help from his teacher, but he still had an uphill battle. So I decided to start a flashcard routine to assist. This solved the original problem. Here is a description of the current routine, and how it has evolved to this point.
It will surprise nobody who has read Teaching Linear Algebra that I started with the thought of some sort of spaced repetition system to maximize his long-term retention with a minimum of effort. I needed to help him with around handwriting, so I wanted to be personally evaluating how he was doing. This seemed simplest with a manual system. I therefore settled on a variation of the Leitner system because that is easy to keep track of by hand.
To make things simple for me to track, I am doing things by powers of 2. Every day we do the whole first pile. Half of the second. A quarter of the third. And so on. (Currently we top out at a 1/256th pile, but are not yet doing any cards from it.) Cards that are done correctly move into the next pile. Those that he get wrong fall into the bad pile, which is the next day's every day pile.
So far, so good. I tried this. Then quickly found that I did an excellent job of sifting through all of the words he knew and getting the ones he didn't know into the bottom pile. But he wasn't learning those. This lead to frustration. Not good.
I then added an extra drill on the pile that he got wrong. At the end of the session, we do a quick drill with just the problem cards. Here is the drill until we get to 3 cards. If he gets the card first try, or gets a card that came all of the way from the bottom since he last got it wrong, it is removed from the drill. If he gets it wrong, I tell him how to do it, and put it back in the pile near the top so he sees it again soon. If he gets it right after a recent reminder, it goes to the bottom to get a chance to come out of the drill.
After we get down to 3 cards, I switch the drill up. If he gets a card wrong I correct him and put it in slot 2. If he gets it right I put it on the bottom. Once he gets all three right, I end the drill for that day.
After I added this final drill on the problem cards, the "not learning" problem disappeared. He began learning, and saw his school performance improve. His spelling tests went from under half the words correct to the 80-100% range. Everyone was happy.
It is worth noting that at the end of grade 1 he took several tests, and we found that he was spelling at a grade 3 level. We have no direct measurement proving it, but I guarantee that he spells even better now.
This happiness lasted until he got used to doing well. Over time we had more piles. In school he was being given more words. I began adding simple arithmetic facts. This meant more and more work. Not fun work. Sometimes he would make a mistake on a card that he had known for a long time. Then he'd get upset. Once he got upset he'd get lots of others wrong. Over the next few days we'd get the cards moving back up the piles, then it would happen again. The flashcard routine became a point of conflict.
Then I had a great idea (which I borrowed from a speech therapist). The idea is that I'd mix a reward activity and flashcards. We'd start on the reward, then do a pile, go back to the reward, then do another pile, go back to the reward, and so on. The specific reward activity that we're using is that I'm reading books to him that are beyond his current reading level (currently The Black Cauldron), but in principle it could be anything. With this shift, the motivation problem completely disappeared. He enjoys the reward. The flashcards are a minor annoyance that gets him the reward. If he goes off track, the reward restores his equilibrium. Intellectually he's happy that he's mastering the material. But the reward is motivation.
With this fix in place, we lasted several months. Then we developed an issue. A couple of words were sufficiently hard that they just stayed in the bad pile every day. So I made a minor tweak. I had been doing his top pile, then his next, then his next, on down. But instead I do his every day pile. Then go into the top pile, next, next, etc. But after each of those groups I try him again on the every day words that he hasn't gotten right yet. Thus he is forced to get his trouble words right 2x per day. This helped him master them and got them moving back up.
With that fix, we lasted until this week. This week we had a problem. His spelling test for this week includes the word embarrassing. (And he can get a bonus for knowing peculiar.) The problem is that this word has enough spelling tricks to get it right that he simply cannot get it in one pass. We tried several times, without success. I therefore have added flashcards like em(barr)assing for which he gets told, "The word 'embarrassing' starts 'em'. Write the 'barr' bit." With these intermediate flashcards he seems to be breaking up learning the whole word into manageable tasks, from which he can learn the word itself. But I've also generated a ton of temporary flashcards, which may become an issue. (I plan on removing those piecemeal ones after he successfully gets them in the every 8 day pile. In a few weeks I'll know how well this is working)
That brings us to the current state of his flashcard routine. He currently has hundreds of spelling words and basic arithmetic facts learned. 373 of them learned sufficiently well that he reviews them less than once per month. But I am sure that I'm not done tweaking. Here are current issues:
It will surprise nobody who has read Teaching Linear Algebra that I started with the thought of some sort of spaced repetition system to maximize his long-term retention with a minimum of effort. I needed to help him with around handwriting, so I wanted to be personally evaluating how he was doing. This seemed simplest with a manual system. I therefore settled on a variation of the Leitner system because that is easy to keep track of by hand.
To make things simple for me to track, I am doing things by powers of 2. Every day we do the whole first pile. Half of the second. A quarter of the third. And so on. (Currently we top out at a 1/256th pile, but are not yet doing any cards from it.) Cards that are done correctly move into the next pile. Those that he get wrong fall into the bad pile, which is the next day's every day pile.
So far, so good. I tried this. Then quickly found that I did an excellent job of sifting through all of the words he knew and getting the ones he didn't know into the bottom pile. But he wasn't learning those. This lead to frustration. Not good.
I then added an extra drill on the pile that he got wrong. At the end of the session, we do a quick drill with just the problem cards. Here is the drill until we get to 3 cards. If he gets the card first try, or gets a card that came all of the way from the bottom since he last got it wrong, it is removed from the drill. If he gets it wrong, I tell him how to do it, and put it back in the pile near the top so he sees it again soon. If he gets it right after a recent reminder, it goes to the bottom to get a chance to come out of the drill.
After we get down to 3 cards, I switch the drill up. If he gets a card wrong I correct him and put it in slot 2. If he gets it right I put it on the bottom. Once he gets all three right, I end the drill for that day.
After I added this final drill on the problem cards, the "not learning" problem disappeared. He began learning, and saw his school performance improve. His spelling tests went from under half the words correct to the 80-100% range. Everyone was happy.
It is worth noting that at the end of grade 1 he took several tests, and we found that he was spelling at a grade 3 level. We have no direct measurement proving it, but I guarantee that he spells even better now.
This happiness lasted until he got used to doing well. Over time we had more piles. In school he was being given more words. I began adding simple arithmetic facts. This meant more and more work. Not fun work. Sometimes he would make a mistake on a card that he had known for a long time. Then he'd get upset. Once he got upset he'd get lots of others wrong. Over the next few days we'd get the cards moving back up the piles, then it would happen again. The flashcard routine became a point of conflict.
Then I had a great idea (which I borrowed from a speech therapist). The idea is that I'd mix a reward activity and flashcards. We'd start on the reward, then do a pile, go back to the reward, then do another pile, go back to the reward, and so on. The specific reward activity that we're using is that I'm reading books to him that are beyond his current reading level (currently The Black Cauldron), but in principle it could be anything. With this shift, the motivation problem completely disappeared. He enjoys the reward. The flashcards are a minor annoyance that gets him the reward. If he goes off track, the reward restores his equilibrium. Intellectually he's happy that he's mastering the material. But the reward is motivation.
With this fix in place, we lasted several months. Then we developed an issue. A couple of words were sufficiently hard that they just stayed in the bad pile every day. So I made a minor tweak. I had been doing his top pile, then his next, then his next, on down. But instead I do his every day pile. Then go into the top pile, next, next, etc. But after each of those groups I try him again on the every day words that he hasn't gotten right yet. Thus he is forced to get his trouble words right 2x per day. This helped him master them and got them moving back up.
With that fix, we lasted until this week. This week we had a problem. His spelling test for this week includes the word embarrassing. (And he can get a bonus for knowing peculiar.) The problem is that this word has enough spelling tricks to get it right that he simply cannot get it in one pass. We tried several times, without success. I therefore have added flashcards like em(barr)assing for which he gets told, "The word 'embarrassing' starts 'em'. Write the 'barr' bit." With these intermediate flashcards he seems to be breaking up learning the whole word into manageable tasks, from which he can learn the word itself. But I've also generated a ton of temporary flashcards, which may become an issue. (I plan on removing those piecemeal ones after he successfully gets them in the every 8 day pile. In a few weeks I'll know how well this is working)
That brings us to the current state of his flashcard routine. He currently has hundreds of spelling words and basic arithmetic facts learned. 373 of them learned sufficiently well that he reviews them less than once per month. But I am sure that I'm not done tweaking. Here are current issues:
- One week is not enough. Every week he is given a new set of words to master. But as anyone who has done spaced repetition knows, a week is not very long to master material. Spaced repetition excels for memorizing a body of data over years, not one week. On most weeks he is given a set of standard words to learn, and a set of words for bonus points. With the bonus words he usually gets over 100% on his tests. But we don't stop, so now he'd do substantially better on last week's test than he actually did last week.
- He's only learning what I know that he needs to. This week I reached out to his teacher and said that I am doing flashcards with him, and looked for feedback on more ways to use them for his benefit. She pointed out a number of things he can improve on, including common words that he has wrong, grammar, poems he is supposed to memorize, and geography that he is supposed to learn. The flashcard routine can help with these issues in time, but I had not been aware that he needed it. Better late than never...
- Work is climbing again. Currently every day I add 2 cards. Plus every week I add a spelling test of unpredictable size (this week 27, of which he already knew one). This is increasing the size of the bottom piles, and the work has been increasing. It is manageable, but I'm keeping my eye on it.
- This takes my time. At the moment that's unavoidable. One of the issues that we're still working on is handwriting, so there needs to be a human evaluation of what he's doing. But still I'm taking an hour per day with this. I think it is an hour well-spent that we both value. However in a couple of years if his sister needs similar help, what then? In the long run I'd love to offload the flashcards to a computer program, but the idea of a reward activity has to be in there. All of the flashcard apps that I've seen assume that doing flashcards is itself a fun activity. That will not work for my son. Maybe I'm being too picky. But I've developed opinions about what works while fine-tuning my son's system. If there is something that fits that, I'd love to find it.
Labels:
children,
education,
flashcards,
Leitner system,
school,
spaced repetition,
spelling
Monday, October 8, 2012
How reliable will the Falcon 9 be?
Let's apply statistics to see, based on current launch data, how reliable we predict that the Falcon 9 will be.
Falcon 9 just had a launch that succeeded despite an engine failure. According to design parameters, it should be able to survive the failure of any two engines. But the flight can be lost if we lose 3+ engines. Exactly how reliable is the Falcon 9 design?
Let me first take a naive approach. To date we've had 4 launches of the Falcon 9, each with 9 engines (that's the 9 in Falcon 9), and have seen one in flight failure. The measured success rate of an engine is therefore 35/36. With that in mind, we can produce the following figures.
For comparison the US Space Shuttle had a failure rate of 2/135 which is about 1.5%.
So SpaceX flights are dangerous compared to most things that we do, but so far seem much better than any previous mode of transport, including the US Space Shuttle. Which was previously the most reliable form of transport into space. (Not the safest though! Soyuz has that record because, unlike the Space Shuttle, they've demonstrated the ability to have passengers survive a catastrophic failure that aborted the mission.)
But is that the end of the story? No!
Suppose that the true failure rate of each individual engine is actually 10%. Then an exactly parallel calculation to the above will find that the failure rate of a rocket launch is 5.3%. That doesn't sound very reliable!
However is it reasonable to think that 10% is a likely failure rate for the rocket? Well suppose that before we had seen any launches that we thought that a 10% failure rate was equally likely as a failure rate of 1/36. Our observation is 1 engine failure out of 36. The odds of that exact observation with a 10% failure rate are 9.0%. The odds of that observation with a failure rate of 1/37 are 37.3%. According to Bayes' theorem, the probabilities that we give to theories after making an observation should be proportional to our initial belief of the probability of that theory times the probability of the given observation under that theory.
That is a mouthful. Let's look at numbers. In this hypothetical scenario our initial belief was a 50% chance of a 10% failure rate, and a 50% chance of a failure rate of 1/36. After observing 36 instances of engines lifting off with 1 failure, the 10% theory has probability proportional to 4.5%, while the 1/36 theory has probability proportional to 18.35%. Thus our updated belief is that the 10% theory has likelihood 4.5/(4.5 + 18.35) = 0.199 = 20%. (Without the intermediate rounding we'd actually be at 0.195.) And the 1/36 theory has likelihood around 80%. Then combining the predictions of the theories with the likelihood assigned to each theory we get an estimated failure rate of 0.053 * 0.195 + 0.0016 * 0.805 = 0.023= 1.16%. Our confidence in the record put up by the Falcon 9 is not as good now!
Please note the following characteristics of this analysis:
Falcon 9 just had a launch that succeeded despite an engine failure. According to design parameters, it should be able to survive the failure of any two engines. But the flight can be lost if we lose 3+ engines. Exactly how reliable is the Falcon 9 design?
Let me first take a naive approach. To date we've had 4 launches of the Falcon 9, each with 9 engines (that's the 9 in Falcon 9), and have seen one in flight failure. The measured success rate of an engine is therefore 35/36. With that in mind, we can produce the following figures.
- Probability of no engine failures: (35/36)**9 * (1 - 35/36)**0 * (9 choose 0) = (35/36)**9 = 77.6%
- Probability of 1 engine failure: (35/36)**8 * (1 - 35/36)**1 * (9 choose 1) = (35/36)**8 * (1/36) * 9 = 20.0%
- Probability of 2 engine failures: (35/36)**7 * (1 - 35/36)**2 * (9 choose 2) = (35/36)**7 * (1/36)**2 * 36 = 1.8%
- Probability of 3+ engine failures: 1 - above probabilities = 0.2% (actually 0.16%)
For comparison the US Space Shuttle had a failure rate of 2/135 which is about 1.5%.
So SpaceX flights are dangerous compared to most things that we do, but so far seem much better than any previous mode of transport, including the US Space Shuttle. Which was previously the most reliable form of transport into space. (Not the safest though! Soyuz has that record because, unlike the Space Shuttle, they've demonstrated the ability to have passengers survive a catastrophic failure that aborted the mission.)
But is that the end of the story? No!
Suppose that the true failure rate of each individual engine is actually 10%. Then an exactly parallel calculation to the above will find that the failure rate of a rocket launch is 5.3%. That doesn't sound very reliable!
However is it reasonable to think that 10% is a likely failure rate for the rocket? Well suppose that before we had seen any launches that we thought that a 10% failure rate was equally likely as a failure rate of 1/36. Our observation is 1 engine failure out of 36. The odds of that exact observation with a 10% failure rate are 9.0%. The odds of that observation with a failure rate of 1/37 are 37.3%. According to Bayes' theorem, the probabilities that we give to theories after making an observation should be proportional to our initial belief of the probability of that theory times the probability of the given observation under that theory.
That is a mouthful. Let's look at numbers. In this hypothetical scenario our initial belief was a 50% chance of a 10% failure rate, and a 50% chance of a failure rate of 1/36. After observing 36 instances of engines lifting off with 1 failure, the 10% theory has probability proportional to 4.5%, while the 1/36 theory has probability proportional to 18.35%. Thus our updated belief is that the 10% theory has likelihood 4.5/(4.5 + 18.35) = 0.199 = 20%. (Without the intermediate rounding we'd actually be at 0.195.) And the 1/36 theory has likelihood around 80%. Then combining the predictions of the theories with the likelihood assigned to each theory we get an estimated failure rate of 0.053 * 0.195 + 0.0016 * 0.805 = 0.023= 1.16%. Our confidence in the record put up by the Falcon 9 is not as good now!
Please note the following characteristics of this analysis:
- Observations do not tell us what reality is, they update our models of reality.
- A wide range of failure probabilities fit the limited observations that we have so far on the Falcon 9.
- With enough data, theories that are far away from the observed average become very unlikely.
Now a curious person might want to know what the odds of failure would be if we included more possible prior theories. I whipped up a quick Perl script to do the calculation for an initial expectation that 0.00%, 0.01%, 0.02%, ..., 99.99%, 100% were all equally likely failure rates a priori. When I run that script I get a probability of 0.0198180199757443, which is an estimated failure rate of about 2%. If you start with different beliefs, you can generate very different specific numbers. For an extreme instance if you believe that SpaceX is constantly improving, so their future engines are likely to be more reliable than their past ones, then ridiculously good numbers become very plausible.
However the bottom line is that we cannot yet, based on the data that we have so far, conclude that we have good evidence that the Falcon 9 actually will put up a better reliability record over its lifetime than previous space vehicles.
Monday, September 17, 2012
A/B Testing vs MAB algorithms - It's complicated
Several months ago, several blog posts appeared on Hacker News comparing A/B testing and multi-armed bandit techniques. If you want to review the posts and the discussion, see 20 lines of code that beat A/B testing every time, then Why multi-armed bandit algorithm is not "better" than A/B testing and finally Why Multi-armed Bandit algorithms are superior to A/B testing (with Math). I participated in those discussions, and ever since then I've been wanting to write up my thoughts once I had them in a compact enough form to do so.
That has taken an unfortunately long time. In fact I've given up on saying everything that I want to say in a compact form, and will try to only say what I think is most important. And even that has wound up less compact than I'd like...
First a disclaimer. Website optimization has been a large part of what I've done in the last decade, and I've been a heavy user of A/B testing. See Effective A/B Testing for a well-regarded tutorial that I did on it several years ago. I have much less experience with multi-armed bandit approaches. I don't believe that I am biased. But if I were, it is clear what my bias would be.
Here is a quick summary of what the rest of this post will attempt to convince you of.
That summary requires a lot of justification. Read on for that.
- If you have actual traffic and are not running tests, start now. I don't actually expand on this fact, but it is trivially true. If you've not started testing, it would be a shock if you can't find at least one 5-10% improvement in your business within 6 months of trying it. Odds are that you'll find several. What is that worth to you?
- A/B testing is an effective optimization methodology.
- A good multi-armed bandit strategy provides another effective optimization methodology. Behind the scenes there is more math, more assumptions, and the possibility of better theoretical characteristics than A/B testing.
- Despite this, if you want to be confident in your statistics, want to be able to do more complex analysis, or have certain business problems, A/B testing likely is a better fit.
- And finally if you want an automated "set and forget" approach, particularly if you need to do continuous optimization, bandit approaches should be considered first.
That summary requires a lot of justification. Read on for that.
Monday, August 8, 2011
Financial stress on the system
I posted this last week on Google+ at https://plus.google.com/u/0/114613808538621741268/posts/13tBzre1nRa. Given market drops since then, I'm reposting it in a more public forum.
How could our next financial crisis start? I was in an interesting discussion about this last night. I'd give better than even odds that it will be ugly before Christmas. Here are some of the possibilities to watch.
- General economic slowdown. The economy seems to be slowing. Plus there are all of the other stresses I'll discuss, all of which are likely to make it slow more. But it seems implausible that a general slowdown will happen suddenly enough to precipitate a full-blown crisis. Though it sure provides stress that could cause something else to happen.
- European sovereign debt issues. We all have heard about Greece. Italy is in the news. But the really scary quantities of debt are in Spain. And Europe has proven to have no political appetite for having meaningful bailouts of one country by another. Either the prospect of the EMU (European Monetary Union) dissolving, or a liquidity crisis caused by a large default (even if it doesn't get called a default) is a likely bet.
- US sovereign debt issue. We dodged the debt ceiling debate. 2 of the 3 rating agencies have left the USA at AAA, but on a negative watch. (Meaning better than even odds of a downgrade within a year.) The third, S&P, based on past guidance is likely to downgrade the USA this month. Though they could do what the other two did. The next test is what the super committee comes up with for the second half of the cuts that are in the debt ceiling deal. That's expected around Thanksgiving. Whether that triggers a crisis depends on what is cut. But a crisis is unlikely. The bigger problem is that people have been backed into a political corner where the US government will be unable to intervene if something else goes south.
- Residential mortgages. The current expectation in the industry is a 3-6% drop next year, with 15% as a worst case scenario. However if the mortgage tax deduction were to be yanked (for instance as part of deficit reduction), much worse changes are possible. (And, of course, another financial crisis would take all bets off of the table.)
- Commercial mortgages. Everyone is aware that commercial mortgages are in a world of pain. However servicers have been very forgiving. They really don't want to force bankruptcies which are in nobody's best interest, so they restructure deals, collect as much blood as they can, then limp along. (A clear case of, "If you owe $100,000, you have a problem. If you owe $100,000,000 the bank has a problem.") However the sector is still under stress. And last month the delinquency rate spiked to the highest rate on record. There could still be trouble there. Particularly if a slowing economy reduces how much money is available to pay back debt.
- Private equity. There is a world of pain coming here as deals struck in 2006 and 2007 (2007 had a lot more) see principal payments come due this year and next. Based on the experience with commercial mortgages, I would expect the banks to be unusually forgiving. But the question is whether they have sufficient liquidity to be forgiving enough. (As much as the banks may want to make deals, they have a real problem if nobody pays them back.) Borders just went bankrupt, how many more companies will follow in the next year?
- Municipal debt. Meredith Whitney has been crying wolf on this for some time, but she has a real point. There are trillions of dollars of debt issued by municipal governments who are under stress. In the first half of this year they got unexpected revenue increases. However a slowing economy could cause them to get back in trouble. And their budget cuts to deal with financial pain could slow the economy. To add salt to the wound, most have large investments in pension plans that is budgeted to get 8% annual growth. Those rates are hard to reach in today's economy, so a lot of that money is in risky asset classes like investments in private equity. It is very likely that those assets will under perform, adding to budget holes. Central Falls, Rhode Island just filed for bankruptcy. How many more will follow in the next year?
- Chinese asset bubble pops. I know little about what is going on there. They do have a big asset bubble. History says that it will pop eventually. Nobody knows when, or what the fallout will be. However my impression is that when it happens, the pain will be mostly felt in China.
- External shocks. You name it. Successful terrorist attack. (The 10'th anniversary of 9/11 is coming up...) Unrest in the Middle East triggers oil spike. An unexpected major earthquake. (For instance the Cascadia fault slips, leaving over 100K dead in Seattle, Vancouver and Victoria.) These are all unlikely, but possible.
How could our next financial crisis start? I was in an interesting discussion about this last night. I'd give better than even odds that it will be ugly before Christmas. Here are some of the possibilities to watch.
- General economic slowdown. The economy seems to be slowing. Plus there are all of the other stresses I'll discuss, all of which are likely to make it slow more. But it seems implausible that a general slowdown will happen suddenly enough to precipitate a full-blown crisis. Though it sure provides stress that could cause something else to happen.
- European sovereign debt issues. We all have heard about Greece. Italy is in the news. But the really scary quantities of debt are in Spain. And Europe has proven to have no political appetite for having meaningful bailouts of one country by another. Either the prospect of the EMU (European Monetary Union) dissolving, or a liquidity crisis caused by a large default (even if it doesn't get called a default) is a likely bet.
- US sovereign debt issue. We dodged the debt ceiling debate. 2 of the 3 rating agencies have left the USA at AAA, but on a negative watch. (Meaning better than even odds of a downgrade within a year.) The third, S&P, based on past guidance is likely to downgrade the USA this month. Though they could do what the other two did. The next test is what the super committee comes up with for the second half of the cuts that are in the debt ceiling deal. That's expected around Thanksgiving. Whether that triggers a crisis depends on what is cut. But a crisis is unlikely. The bigger problem is that people have been backed into a political corner where the US government will be unable to intervene if something else goes south.
- Residential mortgages. The current expectation in the industry is a 3-6% drop next year, with 15% as a worst case scenario. However if the mortgage tax deduction were to be yanked (for instance as part of deficit reduction), much worse changes are possible. (And, of course, another financial crisis would take all bets off of the table.)
- Commercial mortgages. Everyone is aware that commercial mortgages are in a world of pain. However servicers have been very forgiving. They really don't want to force bankruptcies which are in nobody's best interest, so they restructure deals, collect as much blood as they can, then limp along. (A clear case of, "If you owe $100,000, you have a problem. If you owe $100,000,000 the bank has a problem.") However the sector is still under stress. And last month the delinquency rate spiked to the highest rate on record. There could still be trouble there. Particularly if a slowing economy reduces how much money is available to pay back debt.
- Private equity. There is a world of pain coming here as deals struck in 2006 and 2007 (2007 had a lot more) see principal payments come due this year and next. Based on the experience with commercial mortgages, I would expect the banks to be unusually forgiving. But the question is whether they have sufficient liquidity to be forgiving enough. (As much as the banks may want to make deals, they have a real problem if nobody pays them back.) Borders just went bankrupt, how many more companies will follow in the next year?
- Municipal debt. Meredith Whitney has been crying wolf on this for some time, but she has a real point. There are trillions of dollars of debt issued by municipal governments who are under stress. In the first half of this year they got unexpected revenue increases. However a slowing economy could cause them to get back in trouble. And their budget cuts to deal with financial pain could slow the economy. To add salt to the wound, most have large investments in pension plans that is budgeted to get 8% annual growth. Those rates are hard to reach in today's economy, so a lot of that money is in risky asset classes like investments in private equity. It is very likely that those assets will under perform, adding to budget holes. Central Falls, Rhode Island just filed for bankruptcy. How many more will follow in the next year?
- Chinese asset bubble pops. I know little about what is going on there. They do have a big asset bubble. History says that it will pop eventually. Nobody knows when, or what the fallout will be. However my impression is that when it happens, the pain will be mostly felt in China.
- External shocks. You name it. Successful terrorist attack. (The 10'th anniversary of 9/11 is coming up...) Unrest in the Middle East triggers oil spike. An unexpected major earthquake. (For instance the Cascadia fault slips, leaving over 100K dead in Seattle, Vancouver and Victoria.) These are all unlikely, but possible.
Tuesday, July 26, 2011
A worst case for the possible US government default
(This is crossposted from Google+. If you use Google+, you scan find the original here.)
Lots of people are scared about the possibility of a US government default. But I have an unusual reason that few share. And many might think me crazy for worrying about one of the things that I worry about. So I'd like to share why I think what I do before I confess to my worry.
Close to a decade ago I read a fascinating book, Wealth and Democracy by Kevin Phillips. SIt is a large book. It is a detailed book. Most of it is a detailed history of when and where great wealth was created in US history, and its effects on the political process. (In a nutshell the conclusion is that periods of great wealth are periods where the wealthy wield great political influence, and conversely periods where the wealthy have less political control saw great declines of wealth disparities.) Don't try reading it unless you can get into the idea of reading close to 500 pages from a policy wonk going on about things that happened a hundred years ago.
Of course if you are able to read it, there is a lot of interesting material there. Kevin Phillips put a lot of effort in, and knows his stuff. He also has a history of making absurd predictions that work out. For instance in 1970 he wrote The Emerging Republican Majority which made the (then) absurd prediction that white Southerners who reflexively voted against the Republicans because of their role in the Civil War, were about to switch to the Republican party. Nixon followed the strategy that Phillips outlined, and they eventually did. In 1990, when Bush Sr had the highest approval rating of any president ever, Phillips wrote The Politics of Rich and Poor: Wealth and Electorate in the Reagan Aftermath which argued that Bush was vulnerable. Clinton later said that he used it as his election plan.
So what absurd prediction did he make in Wealth and Democracy? Well here is it as best as I remember it. In the last 400 years there have been three other examples of a country that was the sole global superpower. (Spain, the Netherlands, and England.) All three followed a remarkably similar trajectory. All had a major boom that was more based on financial speculation than real production. During this bubble, actual production was outsourced to other countries. This bubble crashed. A couple of years later all three got into a foreign adventure that was supposed to be quick and easy. Said foreign adventure turned out to be much longer and more costly than predicted. Parallel to this there was a "recovery" in which average people didn't really get much ahead. During this period there was even more outsourcing of real production. That ended in a credit crisis some 7-8 years later. Following the credit crisis there was a paper recovery whose effects weren't widely felt. That was not many years later followed by a second financial crisis. After the second financial crisis there was growing outrage in the population. Followed, almost exactly 15 years after the peak, by civil war in 2 of the cases. In the case of England there was civil unrest, but it was overwhelmed by WW I. The outcome was the establishment of a third major political party, Labour.
Incidentally in all three cases one of the countries that production and industrial knowledge was transferred to eventually wound up as the next superpower. Currently the best two candidates for that next superpower role are India and China.
Remember. This was written after the dot com bust, and before the Iraq invasion. Now compare with events of the last decade.
At first I thought, "He did a lot of work, and has a good track record, so I'll not discount it out of hand but I'll keep a lot of skepticism." Ever since then things have played out as predicted, almost like clockwork. I currently see lots of potential triggers for the second financial crisis. So I think that part is likely to play out, and my main hope is that we're more like England with civil unrest than Spain and the Netherlands with an actual civil war. However I have to say that the increasing polarization of the country over the last decade has made civil war much less unthinkable to me as a possibility than it was when I first read the book.
If the next financial crisis is triggered by a political stalemate, that makes me fear that civil war is more likely. Because the event will be one which arrives with already polarized constituencies that each has has a well-articulated story for the crisis that blames the other for not being willing to (cut government|consider any kind of tax). As financial disaster is felt by all, how can tensions fail to escalate along already clearly drawn boundaries? As they escalate, where will it end?
Please, Congress, find a way to compromise. If you fail, trillions of dollars in excess interest payments is the best possible outcome we can hope for. None of us wants to see the worst. Really.
Lots of people are scared about the possibility of a US government default. But I have an unusual reason that few share. And many might think me crazy for worrying about one of the things that I worry about. So I'd like to share why I think what I do before I confess to my worry.
Close to a decade ago I read a fascinating book, Wealth and Democracy by Kevin Phillips. SIt is a large book. It is a detailed book. Most of it is a detailed history of when and where great wealth was created in US history, and its effects on the political process. (In a nutshell the conclusion is that periods of great wealth are periods where the wealthy wield great political influence, and conversely periods where the wealthy have less political control saw great declines of wealth disparities.) Don't try reading it unless you can get into the idea of reading close to 500 pages from a policy wonk going on about things that happened a hundred years ago.
Of course if you are able to read it, there is a lot of interesting material there. Kevin Phillips put a lot of effort in, and knows his stuff. He also has a history of making absurd predictions that work out. For instance in 1970 he wrote The Emerging Republican Majority which made the (then) absurd prediction that white Southerners who reflexively voted against the Republicans because of their role in the Civil War, were about to switch to the Republican party. Nixon followed the strategy that Phillips outlined, and they eventually did. In 1990, when Bush Sr had the highest approval rating of any president ever, Phillips wrote The Politics of Rich and Poor: Wealth and Electorate in the Reagan Aftermath which argued that Bush was vulnerable. Clinton later said that he used it as his election plan.
So what absurd prediction did he make in Wealth and Democracy? Well here is it as best as I remember it. In the last 400 years there have been three other examples of a country that was the sole global superpower. (Spain, the Netherlands, and England.) All three followed a remarkably similar trajectory. All had a major boom that was more based on financial speculation than real production. During this bubble, actual production was outsourced to other countries. This bubble crashed. A couple of years later all three got into a foreign adventure that was supposed to be quick and easy. Said foreign adventure turned out to be much longer and more costly than predicted. Parallel to this there was a "recovery" in which average people didn't really get much ahead. During this period there was even more outsourcing of real production. That ended in a credit crisis some 7-8 years later. Following the credit crisis there was a paper recovery whose effects weren't widely felt. That was not many years later followed by a second financial crisis. After the second financial crisis there was growing outrage in the population. Followed, almost exactly 15 years after the peak, by civil war in 2 of the cases. In the case of England there was civil unrest, but it was overwhelmed by WW I. The outcome was the establishment of a third major political party, Labour.
Incidentally in all three cases one of the countries that production and industrial knowledge was transferred to eventually wound up as the next superpower. Currently the best two candidates for that next superpower role are India and China.
Remember. This was written after the dot com bust, and before the Iraq invasion. Now compare with events of the last decade.
At first I thought, "He did a lot of work, and has a good track record, so I'll not discount it out of hand but I'll keep a lot of skepticism." Ever since then things have played out as predicted, almost like clockwork. I currently see lots of potential triggers for the second financial crisis. So I think that part is likely to play out, and my main hope is that we're more like England with civil unrest than Spain and the Netherlands with an actual civil war. However I have to say that the increasing polarization of the country over the last decade has made civil war much less unthinkable to me as a possibility than it was when I first read the book.
If the next financial crisis is triggered by a political stalemate, that makes me fear that civil war is more likely. Because the event will be one which arrives with already polarized constituencies that each has has a well-articulated story for the crisis that blames the other for not being willing to (cut government|consider any kind of tax). As financial disaster is felt by all, how can tensions fail to escalate along already clearly drawn boundaries? As they escalate, where will it end?
Please, Congress, find a way to compromise. If you fail, trillions of dollars in excess interest payments is the best possible outcome we can hope for. None of us wants to see the worst. Really.
Wednesday, July 20, 2011
I've switched to Google+
I've been interacting on Google+. You can see my new posts there. The content is different, but includes some content which previously would have appeared here, such as this explanation of homophobia..
Therefore if you were following me here, you may want to try to follow me there instead.
Therefore if you were following me here, you may want to try to follow me there instead.
Sunday, February 13, 2011
Finding related items
I've mentioned in the past (not in this blog) that C++ programs can beat databases by many, many orders of magnitude. I learned this while doing a specific fun project. I have permission from the company to discuss that project, and here is a description of that project.
I was working for a contract for ThisNext. Their database has a large number of records of different items, which a large number of people have tagged with arbitrary tags. Think millions of items, millions of distinct tags, and hundreds of millions of times that tags were applied to items. (Often the same tag has been applied to the same item multiple times.) Their idea was that this data could be used to discover which other items are somehow "related" to the current one, and possibly of interest if you find the current one interesting. This is used to populate related items, which is intentionally somewhat whimsical. For instance if you are interested in a pair of silk hotpants then perhaps you might be interested in velvet shorts, lumberjack hotpants or a Bo Peep Jumpsuit.
The suggestions are intentionally somewhat whimsical, and are generated by user activity. The idea behind them is that two items with tags in comment are likely to be related. The more common the tag, the less related they are. For every item they wanted the 30 most related items, sorted in the order of how related they are.
There is an obvious computational problem that you will face. Suppose a popular tag was applied to 30,000 different items. Then for 30,000 items you're going to have to compare to 30,000 other items, which gives you 900,000,000 relationships. This clearly is a scalability problem. So their initial version had implemented this as they just counted the number of tags in common between items, and they skipped any tag that was used over 500 times.
This originally worked fairly well, but as time went by it slowly got worse. The first problem is that their program slowed down as they got more data. When I arrived it was taking a couple of hours to run and was affecting when they could start running backups. The second problem was that if you had an item that had only been tagged once with a popular tag, such as "sunglasses", then no suggestions would be generated for it. They wanted me to try to solve one or both, preferably both, problems.
My first step was to try to figure out the ideal result. And then I was going to try to optimize that. So my first step was to try to find a smooth function that worked over multiple orders of magnitude, that captured the idea that a common tag was less of a relationship than an uncommon tag. The phrase "multiple orders of magnitude" made me think of logarithms, and so I decided that if a tag had been used n times for one item, m times for the other, and s times all told, then log(n + m)/log(s) would be the strength of that relation. The total weight of the connection between two items, I decided, should be the sum of the weights from the shared tags. I decided to break ties in favor of the most tagged other items. (Thinking that this was evidence of popularity, and popular items were probably better.) And any remaining ties in favor of the most recently created item. (Being tagged a lot in a short time is also evidence of popularity.)
Note that this formula is reasonable, but arbitrary. It was chosen not out of any knowledge of what the "right" answer is, but rather as the simplest reasonable formula that looked like it reflected all of the ideas that they wanted to capture.
Next came finding a potentially efficient way to implement it. My idea was that I could do it with two passes. In the first pass I'd compare each item with up to the top 100 items withing a tag. Then I'd sort and find the top 100 related items. In the second pass I'd figure out the full relationship between those items, and sort in the final order. This would prevent the quadratic performance problem that kept popular tags from being analyzed in the first implementation.
If you are familiar with databases then you may wonder how I was planning to pick the top 100 items from each tag, for all tags at once. If I was using Oracle I could just use analytic queries. That feature made it into the SQL 2003 standard, and so support for windowed functions is showing up in other major databases. But I wasn't using a database with support for that. I was using MySQL. However MySQL has an efficient hack for this problem which is described in http://www.xaprb.com/blog/2006/12/02/how-to-number-rows-in-mysql/. (If you're doing a lot of work with MySQL, that hack is worth knowing about.) With that hack it is easy to have a query that will return the top 100 items for each tag.
So I coded up my program as a series of MySQL queries. When it came to actually processing the items there was no way that there was space for the database to handle it in one big query, so I opted for having a temp table into which I stuck ranges of items, which I'd then process in batches. The program ran..but it took 5 days to finish. This obviously was not going to be good enough. I thought about it long and hard, and decided that the best way to optimize would be to try to rewrite it in something other than SQL.
Normally I'd use Perl, but Perl wastes memory fairly liberally, so I was sure that that it couldn't store this dataset in RAM. I am not a C++ programmer, but I noticed is that C++ has vectors. Vectors support the same basic operations that I am used to in Perl. But (unlike Perl arrays) they are memory efficient and very fast. So I talked to the company, got permission to try rewriting this in C++, then went to work.
According to my memory, I used 4 data structures.
Again I wrote my program in 2 passes exactly like before.
For each item the first pass was meant to identify which other items might possibly be the most related. Here it is in pseudocode:
The purpose of the second pass is to find the final list of other items. Here it is in pseudocode:
Remember that the original SQL version took 5 days. The C++ program took under 20 minutes. I got rid of a details file that wasn't actually needed except for debugging, and I stored all of the logarithms that I needed in a vector so I could do an index lookup rather than an expensive floating point calculation. Total run time dropped to under 5 minutes. I left in various assertions that had been a development aid, and called it a job well done. Together with the steps to run queries, extract a dataset, run the program, and do the final upload, end to end I was generating the full result set in something like 15 minutes.
This astounded me. I had good reason to believe that C++ was going to be faster than the database. But 3 orders of magnitude faster? That I did not expect!
I believe that the biggest single difference is the work you need to do to access data.
What is the fastest way to get at an isolated piece of information in the database? You do an index lookup. Meaning that you do a binary search through a tree that then tells you what database row has your information and what page it is to be found on. (Don't forget that walking a binary tree means multiple comparisons, if/then checks, etc.) You then read that page of memory, parse through it to find the row you want, parse the information you want out of that row and return it. Note that the instructions for what exactly to do is likely represented in some sort of byte code that has to be interpreted on the fly.
In my C++ program the equivalent operation is, "Access vector by index." Is there really any surprise that this is several orders of magnitude faster?
There are other differences as well. In the database I was working hard to keep batches small enough that the temporary tables I was throwing around would fit in RAM. In C++ a good deal of the time the next piece of information I wanted would be in L1 or L2 cache inside of the CPU. (The fact that I was walking through data linearly helps here.) MySQL likes to represent numeric types in inefficient ways, while the C++ program uses faster native data types. MySQL had to calculate a lot of logarithms, which was work I mostly avoided in C++. These things all helped.
It is also possible that I had some fixable performance problem in my SQL solution. But in discussing this project I have encountered a number of other people who have had similar experiences. And the general order of magnitude of speedup that I encountered is apparently not particularly rare when going from SQL to C++.
However despite this experience, I remain a big proponent of pushing work to the database and not worrying about performance unless you need to. Why? Because I've pushed databases much harder than most people do. Only a handful of times have I hit their limits. And until I hit those limits, they do a lot to simplify my life, and speed development.
But I am now aware exactly how many orders of magnitude faster I can make things by rewriting in C++. But with a giant corresponding increase in development time and required testing before you know you have correct results.
A final note. This project was fun for me, but what did ThisNext think? They were delighted! I solved their batch processing performance problem. In testing the suggestions produced were clearly better. And based on traffic and revenue figures following the release of my change, I was told that I had increased revenue by at least 10%.
I was working for a contract for ThisNext. Their database has a large number of records of different items, which a large number of people have tagged with arbitrary tags. Think millions of items, millions of distinct tags, and hundreds of millions of times that tags were applied to items. (Often the same tag has been applied to the same item multiple times.) Their idea was that this data could be used to discover which other items are somehow "related" to the current one, and possibly of interest if you find the current one interesting. This is used to populate related items, which is intentionally somewhat whimsical. For instance if you are interested in a pair of silk hotpants then perhaps you might be interested in velvet shorts, lumberjack hotpants or a Bo Peep Jumpsuit.
The suggestions are intentionally somewhat whimsical, and are generated by user activity. The idea behind them is that two items with tags in comment are likely to be related. The more common the tag, the less related they are. For every item they wanted the 30 most related items, sorted in the order of how related they are.
There is an obvious computational problem that you will face. Suppose a popular tag was applied to 30,000 different items. Then for 30,000 items you're going to have to compare to 30,000 other items, which gives you 900,000,000 relationships. This clearly is a scalability problem. So their initial version had implemented this as they just counted the number of tags in common between items, and they skipped any tag that was used over 500 times.
This originally worked fairly well, but as time went by it slowly got worse. The first problem is that their program slowed down as they got more data. When I arrived it was taking a couple of hours to run and was affecting when they could start running backups. The second problem was that if you had an item that had only been tagged once with a popular tag, such as "sunglasses", then no suggestions would be generated for it. They wanted me to try to solve one or both, preferably both, problems.
My first step was to try to figure out the ideal result. And then I was going to try to optimize that. So my first step was to try to find a smooth function that worked over multiple orders of magnitude, that captured the idea that a common tag was less of a relationship than an uncommon tag. The phrase "multiple orders of magnitude" made me think of logarithms, and so I decided that if a tag had been used n times for one item, m times for the other, and s times all told, then log(n + m)/log(s) would be the strength of that relation. The total weight of the connection between two items, I decided, should be the sum of the weights from the shared tags. I decided to break ties in favor of the most tagged other items. (Thinking that this was evidence of popularity, and popular items were probably better.) And any remaining ties in favor of the most recently created item. (Being tagged a lot in a short time is also evidence of popularity.)
Note that this formula is reasonable, but arbitrary. It was chosen not out of any knowledge of what the "right" answer is, but rather as the simplest reasonable formula that looked like it reflected all of the ideas that they wanted to capture.
Next came finding a potentially efficient way to implement it. My idea was that I could do it with two passes. In the first pass I'd compare each item with up to the top 100 items withing a tag. Then I'd sort and find the top 100 related items. In the second pass I'd figure out the full relationship between those items, and sort in the final order. This would prevent the quadratic performance problem that kept popular tags from being analyzed in the first implementation.
If you are familiar with databases then you may wonder how I was planning to pick the top 100 items from each tag, for all tags at once. If I was using Oracle I could just use analytic queries. That feature made it into the SQL 2003 standard, and so support for windowed functions is showing up in other major databases. But I wasn't using a database with support for that. I was using MySQL. However MySQL has an efficient hack for this problem which is described in http://www.xaprb.com/blog/2006/12/02/how-to-number-rows-in-mysql/. (If you're doing a lot of work with MySQL, that hack is worth knowing about.) With that hack it is easy to have a query that will return the top 100 items for each tag.
So I coded up my program as a series of MySQL queries. When it came to actually processing the items there was no way that there was space for the database to handle it in one big query, so I opted for having a temp table into which I stuck ranges of items, which I'd then process in batches. The program ran..but it took 5 days to finish. This obviously was not going to be good enough. I thought about it long and hard, and decided that the best way to optimize would be to try to rewrite it in something other than SQL.
Normally I'd use Perl, but Perl wastes memory fairly liberally, so I was sure that that it couldn't store this dataset in RAM. I am not a C++ programmer, but I noticed is that C++ has vectors. Vectors support the same basic operations that I am used to in Perl. But (unlike Perl arrays) they are memory efficient and very fast. So I talked to the company, got permission to try rewriting this in C++, then went to work.
According to my memory, I used 4 data structures.
- item: A vector of structs. Each struct said how often the item had been tagged, where the item_tags started, and the number of item_tags there were. (The item_id was your position in the vector.)
- item_tag: A vector of structs. Each struct knew the item, the tag, and how many times that tag had been applied to that item. It was sorted by item_id then tag_id.
- tag: A vector of structs. Each struct knew the number of times that tag had been tagged, where the tag_items started, and the number of tag_items there were.
- tag_item: A vector of structs which is exactly like item_tag but this one was sorted by tag and then item.
Again I wrote my program in 2 passes exactly like before.
For each item the first pass was meant to identify which other items might possibly be the most related. Here it is in pseudocode:
foreach tag for that item (found in item_tag, the range is from item_tag):
from tag figure out the range of tag_items I was interested in
(the size of the range was a constant, other_items_limit)
foreach other item found in tag_items:
if item and the other item differ:
save the existence of that relationship in relationships
(a relationship contains the other item id, the weight, and
the popularity of the other tag = how many times it was tagged)
sort relationships by other item id
aggregate relationships by other item id (aggregating weights across tags)
sort relationships by weight, times other item tagged, and other item id
(all descending)
pick top other_items_limit items for the second pass
The purpose of the second pass is to find the final list of other items. Here it is in pseudocode:
foreach other item found in the first pass:
weight := 0
traverse item_tag in parallel for both:
if a common tag is found:
add to the calculated weight
save the relationship
sort relationships by weight, times other item tagged, and other item id
(all descending)
for the top 30 relationships:
print this relationship to a file
Remember that the original SQL version took 5 days. The C++ program took under 20 minutes. I got rid of a details file that wasn't actually needed except for debugging, and I stored all of the logarithms that I needed in a vector so I could do an index lookup rather than an expensive floating point calculation. Total run time dropped to under 5 minutes. I left in various assertions that had been a development aid, and called it a job well done. Together with the steps to run queries, extract a dataset, run the program, and do the final upload, end to end I was generating the full result set in something like 15 minutes.
This astounded me. I had good reason to believe that C++ was going to be faster than the database. But 3 orders of magnitude faster? That I did not expect!
I believe that the biggest single difference is the work you need to do to access data.
What is the fastest way to get at an isolated piece of information in the database? You do an index lookup. Meaning that you do a binary search through a tree that then tells you what database row has your information and what page it is to be found on. (Don't forget that walking a binary tree means multiple comparisons, if/then checks, etc.) You then read that page of memory, parse through it to find the row you want, parse the information you want out of that row and return it. Note that the instructions for what exactly to do is likely represented in some sort of byte code that has to be interpreted on the fly.
In my C++ program the equivalent operation is, "Access vector by index." Is there really any surprise that this is several orders of magnitude faster?
There are other differences as well. In the database I was working hard to keep batches small enough that the temporary tables I was throwing around would fit in RAM. In C++ a good deal of the time the next piece of information I wanted would be in L1 or L2 cache inside of the CPU. (The fact that I was walking through data linearly helps here.) MySQL likes to represent numeric types in inefficient ways, while the C++ program uses faster native data types. MySQL had to calculate a lot of logarithms, which was work I mostly avoided in C++. These things all helped.
It is also possible that I had some fixable performance problem in my SQL solution. But in discussing this project I have encountered a number of other people who have had similar experiences. And the general order of magnitude of speedup that I encountered is apparently not particularly rare when going from SQL to C++.
However despite this experience, I remain a big proponent of pushing work to the database and not worrying about performance unless you need to. Why? Because I've pushed databases much harder than most people do. Only a handful of times have I hit their limits. And until I hit those limits, they do a lot to simplify my life, and speed development.
But I am now aware exactly how many orders of magnitude faster I can make things by rewriting in C++. But with a giant corresponding increase in development time and required testing before you know you have correct results.
A final note. This project was fun for me, but what did ThisNext think? They were delighted! I solved their batch processing performance problem. In testing the suggestions produced were clearly better. And based on traffic and revenue figures following the release of my change, I was told that I had increased revenue by at least 10%.
Labels:
algorithm,
databases,
MySQL,
performance,
related items,
SQL,
tags,
ThisNext
Thursday, February 3, 2011
SQL formatting style
I have written considerably more SQL than most programmers have. And after you've written and maintained several hundred line queries, it becomes obvious that SQL is a programming language like any other. Which means that indentation helps comprehension for all the same reasons that it would in any other language. (I'm continually puzzled by programmers who fail to grasp this, and try to put SQL statements on one line.) But what indentation style should you use?
Here are the principles that lead to my current chosen style.
And here is an example of what this looks like in practice.
If you've never thought about how to format SQL this style is likely to have a few surprises. But try maintaining a lot of SQL that has been formatted in this way, and you should find that it works out quite well.
Here are the principles that lead to my current chosen style.
- Indentation should be in the range of 2-4 spaces. Research indicates that people can read and understand code better with an indent in that range than one which is larger or smaller than that.
- When one thing depends on another, indentation should be increased. This makes the core statements obvious.
- When easy, line things up.
- In any list of things, it should be possible to extend the list by pasting new lines on without changing others. In a language like SQL where you can't have trailing commas, that means putting the comma before the line.
- If you wrap, then operators belong at the front of lines. This makes the logical flow easy to see at a glance. (I got this one from Perl Best Practices.)
- SQL is usually case insensitive, so use underscores in field names to avoid ambiguous parses of things like ExpertsExchange and ExpertSexChange.
And here is an example of what this looks like in practice.
SELECT s.foo
, t.bar
, x.baz
, CASE
WHEN s.blat = 'Hello'
THEN 'Greeting'
WHEN s.blat = 'Goodbye'
THEN 'Farewell'
ELSE 'Unknown'
END as blat_type
FROM table1 s
JOIN table2 t
ON s.some_id = t.some_id
AND s.some_type = 'wanted type here'
JOIN (
SELECT *
FROM yet_another_table yat
WHERE yat.reason = 'Subquery demonstration'
) AS x
ON t.another_id = x.another_id
WHERE s.blat = 'some'
AND t.blot = 'condition'
;
If you've never thought about how to format SQL this style is likely to have a few surprises. But try maintaining a lot of SQL that has been formatted in this way, and you should find that it works out quite well.
Monday, January 3, 2011
How I do Proofs
I was looking through some old papers, and decided to put online some advice that I gave many years ago about how I tackle mathematical proofs. The advice is at How I Do Proofs. This is only lightly edited from the original that I handed out to a course I was teaching in the late 90s.
The course is the one I described at Teaching Linear Algebra. I never got any feedback on my request for improvements, but judging from how well that class did at proofs I believe that the advice was helpful. As one student I asked put it, "Doing proofs is like filling out a shopping list." That is how it should be in a first course where you are learning how to do proofs.
The class where I introduced this advice was fun. I started by handing out printouts of the advice. I then went through the flowchart at a high level. And then I put up a routine theorem from the text, which was the next thing I was supposed to do in the course. And I said, "OK, now that you've all learned how to prove things, you're going to prove this."
You should have seen the shock on their faces! Some started complaining. So I said, "No, seriously. You will all prove this. Just wait and see. It will work."
Then I began. I asked the one person to say what the first step was. A second person to do it. A third person what the next step was. A fourth to do it. And so on. And every time they did something that could be written down, I wrote down exactly what they said. Before long they had proven a result that none of them thought they could prove!
The third of the homework for that night which was on that day's material (if this statement puzzles you, go back to Teaching Linear Algebra and read about the homework strategy in that course) was very heavy on doing proofs. And I can tell that they referred to the handout by one fact. The next day the grader came to me, very puzzled, and asked me to explain this strange piece of reasoning that everyone had used. It was the contrapositive, which is an interesting alternative to proof by contradiction. One of the problems could be solved either with contradiction or the contrapositive, and they all chose the contrapositive because I'd listed it before contradiction.
I considered this a very good thing. One of the major problems that students have with proof by contradiction is that it works too often. Proofs that require a bit of routine algebra or computation can always be rewritten as proof by contradiction. The rewritten proof is correct, but it is unnecessarily confusing and it is better not to have used contradiction. However because contradiction seems to always work, it becomes a sledgehammer that the student always uses. Without noticing that it is sometimes the wrong tool.
The contrapositive doesn't have this drawback. It can solve the same problems as those that "really" needed proof by contradiction. But it doesn't lend itself to solving problems that never needed proof by contradiction. And so students who have been trained to try the contrapositive do not tend to make their simpler proofs over-complicated.
In fact when I drew up the advice I nearly left out contradiction. But it is so widely used that I figured I had to include it. However I tried to subtly discourage its use. And to the best of my knowledge I succeeded, I'm not aware that any students in my class ever used it in any of their proofs.
The course is the one I described at Teaching Linear Algebra. I never got any feedback on my request for improvements, but judging from how well that class did at proofs I believe that the advice was helpful. As one student I asked put it, "Doing proofs is like filling out a shopping list." That is how it should be in a first course where you are learning how to do proofs.
The class where I introduced this advice was fun. I started by handing out printouts of the advice. I then went through the flowchart at a high level. And then I put up a routine theorem from the text, which was the next thing I was supposed to do in the course. And I said, "OK, now that you've all learned how to prove things, you're going to prove this."
You should have seen the shock on their faces! Some started complaining. So I said, "No, seriously. You will all prove this. Just wait and see. It will work."
Then I began. I asked the one person to say what the first step was. A second person to do it. A third person what the next step was. A fourth to do it. And so on. And every time they did something that could be written down, I wrote down exactly what they said. Before long they had proven a result that none of them thought they could prove!
The third of the homework for that night which was on that day's material (if this statement puzzles you, go back to Teaching Linear Algebra and read about the homework strategy in that course) was very heavy on doing proofs. And I can tell that they referred to the handout by one fact. The next day the grader came to me, very puzzled, and asked me to explain this strange piece of reasoning that everyone had used. It was the contrapositive, which is an interesting alternative to proof by contradiction. One of the problems could be solved either with contradiction or the contrapositive, and they all chose the contrapositive because I'd listed it before contradiction.
I considered this a very good thing. One of the major problems that students have with proof by contradiction is that it works too often. Proofs that require a bit of routine algebra or computation can always be rewritten as proof by contradiction. The rewritten proof is correct, but it is unnecessarily confusing and it is better not to have used contradiction. However because contradiction seems to always work, it becomes a sledgehammer that the student always uses. Without noticing that it is sometimes the wrong tool.
The contrapositive doesn't have this drawback. It can solve the same problems as those that "really" needed proof by contradiction. But it doesn't lend itself to solving problems that never needed proof by contradiction. And so students who have been trained to try the contrapositive do not tend to make their simpler proofs over-complicated.
In fact when I drew up the advice I nearly left out contradiction. But it is so widely used that I figured I had to include it. However I tried to subtly discourage its use. And to the best of my knowledge I succeeded, I'm not aware that any students in my class ever used it in any of their proofs.
Labels:
contrapositive,
linear algebra,
proof by contradiction,
proofs,
teaching
Monday, November 15, 2010
Bluer than blue
On my profile I have answered the question Something I can't find using Google with The lyrics to the version of "Bluer than blue" we sang in highschool jazz. It isn't that you can't find lyrics to the song. The problem is that there are many songs out there with that name, that have unrelated lyrics.
Well my question has been answered. Someone who was in my high school jazz choir found my profile, emailed me, and emailed her memory of the lyrics. Courtesy of Lisa Schmidt, here they are:
Bluer than Blue
Once I thought the world was ours to share
Always knew there'd be someone to care
Now I realize
Promises were lies
Wish I'd find the one who really loves me
I'm feelin' bluer than blue
I'm all alone and feeling bluer than blue
My baby left me, now I'm cryin' like a fool
I want somebody who will love me
Oh well I'm bluer than blue
I'm all alone and need somebody that's true
Looking around to see if you'll come back to me
I want somebody who cares
Guess I've learned my lesson / Broken all the rules
Well, you know I try to recall
One thing I am certain / Love's a game of fools
You're never ready when loneliness calls
Bluer than blue
I'm all alone and need somebody that's true
Hanging around and hoping you'll come back to me
I want somebody who cares
[Here there was a vocal interlude, where we all went "doo, doo do dooo, doo wop etc."]
Guess I learned my lesson / When we started kissin'
Well, you know I try to recall
One thing I am certain / love's a word for hurtin'
You're never ready when loneliness starts calling [key change]
Bluer than blue
I'm all alone and need somebody that's true
Looking around to see if you'll come back to me
I want somebody who will love me, come on baby love me
Can't you see I'm lonely?
I want somebody who's gonna care [she's bluer than bluuuuuuuue]
Thank you Lisa!
Now a bit of reminiscing. When I was in choir, Canada did something really good. What they did is arrange for a variety of regional competitions, some of which would qualify you for the nationals. In those competitions each choir had 15 minutes to do a prepared set. Then one of the judges would come up and work with the choir for 15 minutes. Then the choir would go back stage and be given a workshop. I have no idea if they still do it, but I can't praise it highly enough.
The point of this was not the competition, though that was serious, but rather that it taught the teachers how to teach better. This quickly made everyone better. A lot of the choirs were simply amazing. I had the fortune to be at Esquimalt when the choir teacher, Eileen Cooper, began really going through this system and got mentored by the choir teacher at Argyle. In the time we were there, we went from a mediocre choir using instrumental mikes and with no idea how to really hold them to being among the top high school jazz choirs in the country. To this day I believe that I was only tolerated because it was so hard to find baritones who were willing to be part of the choir.
But be that as it may, most of my positive memories of high school involve choir. And as a parent it has been wonderful that I am able and willing to sing for my kids.
Well my question has been answered. Someone who was in my high school jazz choir found my profile, emailed me, and emailed her memory of the lyrics. Courtesy of Lisa Schmidt, here they are:
Bluer than Blue
Once I thought the world was ours to share
Always knew there'd be someone to care
Now I realize
Promises were lies
Wish I'd find the one who really loves me
I'm feelin' bluer than blue
I'm all alone and feeling bluer than blue
My baby left me, now I'm cryin' like a fool
I want somebody who will love me
Oh well I'm bluer than blue
I'm all alone and need somebody that's true
Looking around to see if you'll come back to me
I want somebody who cares
Guess I've learned my lesson / Broken all the rules
Well, you know I try to recall
One thing I am certain / Love's a game of fools
You're never ready when loneliness calls
Bluer than blue
I'm all alone and need somebody that's true
Hanging around and hoping you'll come back to me
I want somebody who cares
[Here there was a vocal interlude, where we all went "doo, doo do dooo, doo wop etc."]
Guess I learned my lesson / When we started kissin'
Well, you know I try to recall
One thing I am certain / love's a word for hurtin'
You're never ready when loneliness starts calling [key change]
Bluer than blue
I'm all alone and need somebody that's true
Looking around to see if you'll come back to me
I want somebody who will love me, come on baby love me
Can't you see I'm lonely?
I want somebody who's gonna care [she's bluer than bluuuuuuuue]
Thank you Lisa!
Now a bit of reminiscing. When I was in choir, Canada did something really good. What they did is arrange for a variety of regional competitions, some of which would qualify you for the nationals. In those competitions each choir had 15 minutes to do a prepared set. Then one of the judges would come up and work with the choir for 15 minutes. Then the choir would go back stage and be given a workshop. I have no idea if they still do it, but I can't praise it highly enough.
The point of this was not the competition, though that was serious, but rather that it taught the teachers how to teach better. This quickly made everyone better. A lot of the choirs were simply amazing. I had the fortune to be at Esquimalt when the choir teacher, Eileen Cooper, began really going through this system and got mentored by the choir teacher at Argyle. In the time we were there, we went from a mediocre choir using instrumental mikes and with no idea how to really hold them to being among the top high school jazz choirs in the country. To this day I believe that I was only tolerated because it was so hard to find baritones who were willing to be part of the choir.
But be that as it may, most of my positive memories of high school involve choir. And as a parent it has been wonderful that I am able and willing to sing for my kids.
Labels:
bluer than blue,
competition,
Esquimalt,
High school,
jazz,
lyrics
Tuesday, October 26, 2010
Practice poker online for free
I enjoy playing poker. I'm not that good, but I do OK in small home games and I enjoy myself.
Recently I've been having a lot of fun playing online for free. How, you ask? Well go to http://www.fulltiltpoker.com/, sign up for free, get a free account, and start playing with play chips. Work your way up to the level you've got skills for, then have fun.
Now everyone's first reaction when I tell them this is that people play really stupidly when they are playing online for free. It isn't worth trying.
But this is not necessarily so! When you play on Full Tilt you can reload to 1000 play chips every 5 minutes. And at the bottom tables people play like complete donks because they can with impunity. However there are a lot of different buy-ins. As you move up, the skill level increases. So you get a gradation of skill from ridiculously soft at 1/2 up to a boss level of 50K/100K. At the latter level typical buy-ins are 4,000,000. The game I just looked at had 2 players, each with about 25,000,000 in front of them. You can't play that game unless you've managed to earn your way to that level. Anyone who does that has legitimately beaten a *lot* of other players, and put a lot of time in. Do you think those people play poorly?
If you've got basic skills (easily acquired), it isn't hard to beat the easy "donk" levels of free poker online. Then you can move up to more fun levels that require more skill. And once you can play poker well? Well, then you can buy in with some confidence about your skills (at least for low level play) and enjoy both the game and the money you're making at it.
Here is a data point. My sister read a couple of books, went online, got a free account, worked her bankroll up to a few hundred thousand and was playing at the 100/200 level. Then she began playing for money, entered the WPS, and look what happened.
So if you find poker fun, and want to practice, you can online for free. The only thing you'll lose is your time. Which isn't a cost if it is fun for you.
Recently I've been having a lot of fun playing online for free. How, you ask? Well go to http://www.fulltiltpoker.com/, sign up for free, get a free account, and start playing with play chips. Work your way up to the level you've got skills for, then have fun.
Now everyone's first reaction when I tell them this is that people play really stupidly when they are playing online for free. It isn't worth trying.
But this is not necessarily so! When you play on Full Tilt you can reload to 1000 play chips every 5 minutes. And at the bottom tables people play like complete donks because they can with impunity. However there are a lot of different buy-ins. As you move up, the skill level increases. So you get a gradation of skill from ridiculously soft at 1/2 up to a boss level of 50K/100K. At the latter level typical buy-ins are 4,000,000. The game I just looked at had 2 players, each with about 25,000,000 in front of them. You can't play that game unless you've managed to earn your way to that level. Anyone who does that has legitimately beaten a *lot* of other players, and put a lot of time in. Do you think those people play poorly?
If you've got basic skills (easily acquired), it isn't hard to beat the easy "donk" levels of free poker online. Then you can move up to more fun levels that require more skill. And once you can play poker well? Well, then you can buy in with some confidence about your skills (at least for low level play) and enjoy both the game and the money you're making at it.
Here is a data point. My sister read a couple of books, went online, got a free account, worked her bankroll up to a few hundred thousand and was playing at the 100/200 level. Then she began playing for money, entered the WPS, and look what happened.
So if you find poker fun, and want to practice, you can online for free. The only thing you'll lose is your time. Which isn't a cost if it is fun for you.
Wednesday, August 25, 2010
Piaget Water Level Test
Many years ago in a dentist's office I read an interesting article in a magazine. It talked about how there were questions that were specific to certain adult genders. In particular until puberty there was no measurable difference in performance on the question, but after puberty there was a large difference. As examples they offered a verbal question that women do well on, and a question where they drew two cups, one tilted, and asked people to fill in the water level if both were half full. (Artistry not required.) Men do well on this. Most women don't. And education doesn't matter, women who graduate college do worse than men who drop out of high-school.
I botched the verbal one due to something that looked silly to me, got the male one, and dismissed the article as garbage. Then a month later it came up in a conversation with my girlfriend, and she got the men's question wrong. I couldn't remember the one that women do well on. This made me curious, so I asked my mother the same question and she got it wrong. I can call my mother many things, but unintelligent is very much not among them. As an example, when she was at Stanford in the 50s they gave her a battery of ability tests. The only one she was not in the top 1% on was manual dexterity.
(Side note for married men. Do not rush to give your wife this test. I've started many marital fights that way, and I have never tracked down a question that women do better than men on. I know there is one, and I know that at 19 I couldn't do it, but I don't remember what it was and have never encountered another.)
Since then I've learned that this question is called the Piaget Water Level Test. The background on this is that Piaget found that children typically gain specific mental abilities at specific ages that are tied to specific growth spurts. And so he collected examples. For instance there is a specific age at which children learn that people who were not present would not have seen what they did, and another at which children learn that when you pour water into a tall thin glass you don't have more than you used to. And as he collected examples, he eventually offered the water level test. Which, oddly, men gain the ability to do during our last growth spurt, puberty, and most women never do.
I've seen various estimates for how well people do on it. One was that 90% of men can do the task, and only 30% of women. That seems to be a good fit with my experience.
An interesting side note. It is widely noted that there is a large gender imbalance within programming. But I've found through experience that programmers I know, whether male or female, have a 100% success rate on this question. I have no idea what to make of this tidbit, but I find it interesting.
Incidentally for those wondering what the answer to the question is, the water level is horizontal to the ground. The most common answer among women I've asked is to draw the water level parallel to the bottom of the cup. The second most common answer from women is to realize that there is a trick, and to draw the water level tilted twice as much as the cup. When the correct answer is pointed out, women recognize it as very obvious. I imagine that their feeling is much like how I felt after botching the verbal question that I blanked out of my memory.
BTW if anyone knows of any question with the reverse gender characteristics, I've been looking for it for over 20 years. It is frustrating - I know that at least one such question exists, but I've never found it.
I botched the verbal one due to something that looked silly to me, got the male one, and dismissed the article as garbage. Then a month later it came up in a conversation with my girlfriend, and she got the men's question wrong. I couldn't remember the one that women do well on. This made me curious, so I asked my mother the same question and she got it wrong. I can call my mother many things, but unintelligent is very much not among them. As an example, when she was at Stanford in the 50s they gave her a battery of ability tests. The only one she was not in the top 1% on was manual dexterity.
(Side note for married men. Do not rush to give your wife this test. I've started many marital fights that way, and I have never tracked down a question that women do better than men on. I know there is one, and I know that at 19 I couldn't do it, but I don't remember what it was and have never encountered another.)
Since then I've learned that this question is called the Piaget Water Level Test. The background on this is that Piaget found that children typically gain specific mental abilities at specific ages that are tied to specific growth spurts. And so he collected examples. For instance there is a specific age at which children learn that people who were not present would not have seen what they did, and another at which children learn that when you pour water into a tall thin glass you don't have more than you used to. And as he collected examples, he eventually offered the water level test. Which, oddly, men gain the ability to do during our last growth spurt, puberty, and most women never do.
I've seen various estimates for how well people do on it. One was that 90% of men can do the task, and only 30% of women. That seems to be a good fit with my experience.
An interesting side note. It is widely noted that there is a large gender imbalance within programming. But I've found through experience that programmers I know, whether male or female, have a 100% success rate on this question. I have no idea what to make of this tidbit, but I find it interesting.
Incidentally for those wondering what the answer to the question is, the water level is horizontal to the ground. The most common answer among women I've asked is to draw the water level parallel to the bottom of the cup. The second most common answer from women is to realize that there is a trick, and to draw the water level tilted twice as much as the cup. When the correct answer is pointed out, women recognize it as very obvious. I imagine that their feeling is much like how I felt after botching the verbal question that I blanked out of my memory.
BTW if anyone knows of any question with the reverse gender characteristics, I've been looking for it for over 20 years. It is frustrating - I know that at least one such question exists, but I've never found it.
Labels:
development,
gender differences,
piaget,
puberty,
water level test
Thursday, August 19, 2010
Analysis vs Algebra predicts eating corn?
I like learning about odd connections between disparate things. This probably is the oddest example that I know.
Broadly speaking, mathematicians can be divided into those who like analysis, and those who like algebra. The distinction between the two types runs throughout math. Even those who work in areas that are far from analysis or algebra are very aware of the difference between them, and usually are very clear on which their preference is. I'll delve into this in more depth soon, but for now let's just take it for granted that this is a well-known distinction, and it has meaning for mathematicians.
Back when I was in grad school there was a department lunch with corn on the cob. Partway through the meal one of the analysts looked around the room and remarked, "That's odd, all of the analysts are eating corn one way and the algebraists are eating corn another!" Everyone looked around. In fact everyone was eating the corn in one of two ways. One way was to munch over the length of the corn in a straight line, back up, turn slightly, and do another row across. Kind of like how an old typewriter goes. The other way was to go around in a spiral. All of the analysts were eating in spirals, and the algebraists in rows.
There were a number of mathematicians present whose fields of study didn't make it clear whether they were on the analysis or algebra side of things. We went around and asked, and in every case the way they ate corn matched their preference. Since then I've made a point of amusing myself by asking mathematicians I meet whether they prefer algebra or analysis, and then predicting which way they will eat corn. I'm probably up to 40 or so by now, and in every case but one I've been able to correctly predict how they eat corn. The one exception was a logician who claimed to be exactly on the fence between the two. When I explained the corn thing to him he looked surprised, and said that he had an unusual way of eating corn. He went in loose spirals! In other words he truly was a perfect combination of algebra and analysis!
If you have even a passing familiarity of probability, it is clear that despite how unbelievable it initially is that the type of mathematics you prefer is connected to how you eat corn, it is pretty much certain that there actually is a very strong connection. If you believe, as I do, that this difference is connected to how we think about other things, then there must be some odd connection between how we like to understand the world and how we eat corn. Why is another matter.
How do I explain the distinction between algebra and analysis? Well the best way to understand it is to ask you to study advanced mathematics. You will have to take many courses with the word "algebra" in the name, and others with "analysis" in the name. By the time you're done you'll have experienced the difference, and you'll be clear on which you prefer. Odds are you won't do that, but that is the most reliable way to come to understand it.
If I have to wave my hands and explain it, I would explain it like this. In algebra there are sequences of operations which have proven to be important and effective in one circumstance. Algebraists try to reuse these operations in different contexts in the hopes that what proved effective in one situation will be effective again. By contrast an analyst is likely to form an idiosyncratic mental model of specific problems. Based on that mental model you have intuitions that let you carry out long chains of calculations that are, in principle, obviously going to lead to the right thing. Typically your intuition is correct to within a constant factor, and you're only interested in some sort of limiting behavior so that is fine.
If you don't know any advanced math, the odds are about equal that my explanation is going to mislead you as to give you an idea what I am talking about. You'd be better off figuring out your preference by looking at how you eat corn. That said, the distinction carries through into other subjects that I've learned about. But not in a clear and obvious way.
For instance I've noticed the difference cropping up in programming. The distinction is often hard to explain. There are a wide variety of programming techniques, and most programmers have only really learned a few. Some of those techniques appeal to analysts, and others to algebraists. But if you've only been exposed to techniques that are a good fit for one, then how do you know which you'd prefer? Worse yet, when two programmers talk and have different experience bases, how can they tell whether their natural intellectual tastes are similar or different?
Let me give some examples. Upon my first encounter it was clear to me that object oriented programming is something that appeals to algebraists. So if you're a programmer and found Design Patterns: Elements of Reusable Object-Oriented Software to be a revelation, it is highly likely that you lean towards algebra and eat your corn in neat rows. Going the other way, if the techniques described in On Lisp appeal, then you might be on the analytic side of the fence and eat your corn in spirals. This is particularly true if you found yourself agreeing with Paul Graham's thoughts in Why Arc Isn't Especially Object-Oriented. There was a period that I thought that the programming division might be as simple as functional versus object oriented. Then I encountered monads, and I learned that there were functional programmers who clearly were algebraists. (I know someone who got his PhD studying Haskell's type system. My prediction that he ate corn in rows was correct.) Going the other way I wouldn't be surprised that people who love what they can do with template metaprogramming in C++ lean towards analysis and eating corn in spirals. (I haven't tested the last guess at all, so take it with a grain of salt.)
Going out on a limb, I wouldn't be surprised to find out that where people fall in the emacs/vi debate is correlated with how they eat corn. I wouldn't predict a very strong correlation, but I'd expect that emacs is likely to appeal to people who would like algebra, and vi to people who like analysis.
And now to wrap up, why would how we eat corn say something how we think? Here is what I think.
When you pick up a piece of corn on the cob, you have two cues for how to eat it. The first is that the corn is laid out in very nice rows. How can you not follow the lines that are laid out for you? The other is that as you eat, your teeth scrape down the corn. If you twist your wrist, you'll eat more efficiently. Why would someone want to eat inefficiently?
My best guess is that the cue you notice and follow reflects a natural tendency about how you tend to think in general. And this tendency is tied to such things as what kind of math you prefer or what programming techniques would prove interesting for you.
Broadly speaking, mathematicians can be divided into those who like analysis, and those who like algebra. The distinction between the two types runs throughout math. Even those who work in areas that are far from analysis or algebra are very aware of the difference between them, and usually are very clear on which their preference is. I'll delve into this in more depth soon, but for now let's just take it for granted that this is a well-known distinction, and it has meaning for mathematicians.
Back when I was in grad school there was a department lunch with corn on the cob. Partway through the meal one of the analysts looked around the room and remarked, "That's odd, all of the analysts are eating corn one way and the algebraists are eating corn another!" Everyone looked around. In fact everyone was eating the corn in one of two ways. One way was to munch over the length of the corn in a straight line, back up, turn slightly, and do another row across. Kind of like how an old typewriter goes. The other way was to go around in a spiral. All of the analysts were eating in spirals, and the algebraists in rows.
There were a number of mathematicians present whose fields of study didn't make it clear whether they were on the analysis or algebra side of things. We went around and asked, and in every case the way they ate corn matched their preference. Since then I've made a point of amusing myself by asking mathematicians I meet whether they prefer algebra or analysis, and then predicting which way they will eat corn. I'm probably up to 40 or so by now, and in every case but one I've been able to correctly predict how they eat corn. The one exception was a logician who claimed to be exactly on the fence between the two. When I explained the corn thing to him he looked surprised, and said that he had an unusual way of eating corn. He went in loose spirals! In other words he truly was a perfect combination of algebra and analysis!
If you have even a passing familiarity of probability, it is clear that despite how unbelievable it initially is that the type of mathematics you prefer is connected to how you eat corn, it is pretty much certain that there actually is a very strong connection. If you believe, as I do, that this difference is connected to how we think about other things, then there must be some odd connection between how we like to understand the world and how we eat corn. Why is another matter.
How do I explain the distinction between algebra and analysis? Well the best way to understand it is to ask you to study advanced mathematics. You will have to take many courses with the word "algebra" in the name, and others with "analysis" in the name. By the time you're done you'll have experienced the difference, and you'll be clear on which you prefer. Odds are you won't do that, but that is the most reliable way to come to understand it.
If I have to wave my hands and explain it, I would explain it like this. In algebra there are sequences of operations which have proven to be important and effective in one circumstance. Algebraists try to reuse these operations in different contexts in the hopes that what proved effective in one situation will be effective again. By contrast an analyst is likely to form an idiosyncratic mental model of specific problems. Based on that mental model you have intuitions that let you carry out long chains of calculations that are, in principle, obviously going to lead to the right thing. Typically your intuition is correct to within a constant factor, and you're only interested in some sort of limiting behavior so that is fine.
If you don't know any advanced math, the odds are about equal that my explanation is going to mislead you as to give you an idea what I am talking about. You'd be better off figuring out your preference by looking at how you eat corn. That said, the distinction carries through into other subjects that I've learned about. But not in a clear and obvious way.
For instance I've noticed the difference cropping up in programming. The distinction is often hard to explain. There are a wide variety of programming techniques, and most programmers have only really learned a few. Some of those techniques appeal to analysts, and others to algebraists. But if you've only been exposed to techniques that are a good fit for one, then how do you know which you'd prefer? Worse yet, when two programmers talk and have different experience bases, how can they tell whether their natural intellectual tastes are similar or different?
Let me give some examples. Upon my first encounter it was clear to me that object oriented programming is something that appeals to algebraists. So if you're a programmer and found Design Patterns: Elements of Reusable Object-Oriented Software to be a revelation, it is highly likely that you lean towards algebra and eat your corn in neat rows. Going the other way, if the techniques described in On Lisp appeal, then you might be on the analytic side of the fence and eat your corn in spirals. This is particularly true if you found yourself agreeing with Paul Graham's thoughts in Why Arc Isn't Especially Object-Oriented. There was a period that I thought that the programming division might be as simple as functional versus object oriented. Then I encountered monads, and I learned that there were functional programmers who clearly were algebraists. (I know someone who got his PhD studying Haskell's type system. My prediction that he ate corn in rows was correct.) Going the other way I wouldn't be surprised that people who love what they can do with template metaprogramming in C++ lean towards analysis and eating corn in spirals. (I haven't tested the last guess at all, so take it with a grain of salt.)
Going out on a limb, I wouldn't be surprised to find out that where people fall in the emacs/vi debate is correlated with how they eat corn. I wouldn't predict a very strong correlation, but I'd expect that emacs is likely to appeal to people who would like algebra, and vi to people who like analysis.
And now to wrap up, why would how we eat corn say something how we think? Here is what I think.
When you pick up a piece of corn on the cob, you have two cues for how to eat it. The first is that the corn is laid out in very nice rows. How can you not follow the lines that are laid out for you? The other is that as you eat, your teeth scrape down the corn. If you twist your wrist, you'll eat more efficiently. Why would someone want to eat inefficiently?
My best guess is that the cue you notice and follow reflects a natural tendency about how you tend to think in general. And this tendency is tied to such things as what kind of math you prefer or what programming techniques would prove interesting for you.
Labels:
analysis,
corn on the cob,
eating corn,
lalgebra,
programming techniques
Tuesday, August 3, 2010
How did pterosaurs get so big?
The pterosaurs got going something like 230 million years ago. They died out with the dinosaurs 65 million years ago. Over their history they came in all sizes, from Rhamphorhynchus who was the size of a sparrow to Quetzalcoatlus with a wingspan of variously estimated as being 30-40 feet. Spread out it was a similar size to a t-rex, and on the ground its estimated height was close to a giraffe's. The largest bird ever, Argentavis was much, much smaller than that.
However the birds arose about 150 million years ago. Feathers were a big advantage in flight, and over time the birds took over a lot of what the pterosaurs were doing. But the pterosaurs did not go away. Instead they wound up being in niches for very big flying animals. This is competition through specialization, which I talked about some time ago when I discussed the Neanderthals.
This coexistence provides evidence that birds were generally better fliers, but there was a niche for very big fliers that the pterosaurs were better at. The question I'm curious about is why the pterosaurs were better at being big fliers than birds were.
I have a theory. But before I can explain it I need to provide some background.
Wings have evolved in vertebrates three times in pterosaurs, birds, and bats. All three started with the basic vertebrate limb structure and found different ways of constructing a wing out of it. In both bats and birds the arm bones form part of the wing. Now an important fact about vertebrate bones is that different bones grow at different rates as you grow. In particular arm bones start off shorter and catch up later. The result is that in birds and bats, babies have the wrong proportions for their wings to be useful. Therefore baby birds and bats can't learn to fly until they have achieved a significant fraction of their full size.
Pterosaurs were different. Their wings were entirely constructed from wrist and hand bones. (Fully half the wing was supported by an elongated 4th finger.) Hand bones stay in proportion your whole life. Comparisons of fossils of pterosaurs at different ages in the same species verifies that their wings always had good proportions for flight. Furthermore we have fossils from baby pterosaurs that died miles out at sea, which is direct evidence that they flew young.
What does this have to do with the eventual size of the animals? Well birds cannot learn to fly until they are near full growth. Which means that they need intensive care from their parents until they reach that growth. This care is a significant fact of life for bird species, and is why most types of birds have both parents providing care. Unlike most mammals where the mother is generally capable of taking care of young on her own. The larger the bird is, the harder this care is to provide.
By contrast pterosaurs were probably able to take care of themselves at a much younger age, and smaller size. Which means that they were free to grow for a lot longer, to a lot larger size, without unduly taxing their parents. (In truth we don't have any data indicating how much or little parental care baby pterosaurs got. But I suspect it was less than birds get.) And, I believe, that is why they were able to get so much larger than birds.
Random trivia I came across in preparing this post. The reason bats can't fly during the day is that their wings are vulnerable to sunburn. There is evidence that pterosaurs had a protective layer so they didn't have this issue. Also birds have stiffer wings than bats do, which provides better lift and less maneuverability. Pterosaurs had more joints in their wings than birds do, but didn't have finger bones inside of the structure of their wings like bats, which suggests to me that their wings would have been somewhere between.
And my whole train of thought was started by watching National Geographic - Sky Monsters. If you're interested in pterosaurs, it is a worthwhile video.
However the birds arose about 150 million years ago. Feathers were a big advantage in flight, and over time the birds took over a lot of what the pterosaurs were doing. But the pterosaurs did not go away. Instead they wound up being in niches for very big flying animals. This is competition through specialization, which I talked about some time ago when I discussed the Neanderthals.
This coexistence provides evidence that birds were generally better fliers, but there was a niche for very big fliers that the pterosaurs were better at. The question I'm curious about is why the pterosaurs were better at being big fliers than birds were.
I have a theory. But before I can explain it I need to provide some background.
Wings have evolved in vertebrates three times in pterosaurs, birds, and bats. All three started with the basic vertebrate limb structure and found different ways of constructing a wing out of it. In both bats and birds the arm bones form part of the wing. Now an important fact about vertebrate bones is that different bones grow at different rates as you grow. In particular arm bones start off shorter and catch up later. The result is that in birds and bats, babies have the wrong proportions for their wings to be useful. Therefore baby birds and bats can't learn to fly until they have achieved a significant fraction of their full size.
Pterosaurs were different. Their wings were entirely constructed from wrist and hand bones. (Fully half the wing was supported by an elongated 4th finger.) Hand bones stay in proportion your whole life. Comparisons of fossils of pterosaurs at different ages in the same species verifies that their wings always had good proportions for flight. Furthermore we have fossils from baby pterosaurs that died miles out at sea, which is direct evidence that they flew young.
What does this have to do with the eventual size of the animals? Well birds cannot learn to fly until they are near full growth. Which means that they need intensive care from their parents until they reach that growth. This care is a significant fact of life for bird species, and is why most types of birds have both parents providing care. Unlike most mammals where the mother is generally capable of taking care of young on her own. The larger the bird is, the harder this care is to provide.
By contrast pterosaurs were probably able to take care of themselves at a much younger age, and smaller size. Which means that they were free to grow for a lot longer, to a lot larger size, without unduly taxing their parents. (In truth we don't have any data indicating how much or little parental care baby pterosaurs got. But I suspect it was less than birds get.) And, I believe, that is why they were able to get so much larger than birds.
Random trivia I came across in preparing this post. The reason bats can't fly during the day is that their wings are vulnerable to sunburn. There is evidence that pterosaurs had a protective layer so they didn't have this issue. Also birds have stiffer wings than bats do, which provides better lift and less maneuverability. Pterosaurs had more joints in their wings than birds do, but didn't have finger bones inside of the structure of their wings like bats, which suggests to me that their wings would have been somewhere between.
And my whole train of thought was started by watching National Geographic - Sky Monsters. If you're interested in pterosaurs, it is a worthwhile video.
Labels:
bats,
birds,
evolution,
flight,
pterosaurs,
speculation
Tuesday, July 6, 2010
Dozenal glyphs - a modest proposal
A proposal that shows up every so often is to use dozenal arithmetic. The idea is simple, in base 12 more fractions work out very conveniently. You can easily divide things 3 and 4 ways. This is frequently convenient, and is why we often sell things in dozens. It is also why the much maligned Imperial system sneaks factors of 3 (3 tsp in a tbsp) and 12 (inches in a foot) in various places. And there are numerous minor benefits, such as the fact that the multiplication table becomes significantly simpler and therefore easier to learn.
When the French created the metric system, they based it on factors of 10 everywhere. They even went so far as to try to measure angles in gradians (a quarter circle had 100), and to use decimal time. The world as a whole rejected decimal angles and times, but has adopted decimal metric everywhere else. Which is very convenient for scientists, but is hard to divide into thirds and somewhat inconvenient for quarters. Furthermore the persistence of both systems results in occasional annoyances like the fact that in daily life we usually prefer to measure speed in km/h, but for energy and power calculations the units only work out properly if you measure in m/s. Resulting in an annoying factor of 3.6 that shows up converting between them.
The other day I read An Argument for Dozenalism that made many of these arguments. Nothing new. However it made the interesting point that ideally a dozenal arithmetic would have its own set of glyphs. It suggested that 0 and 1 could be kept, but everything else should be changed. In my opinion that argument is correct, it is very confusing if 21 sometimes is 5*5 and sometimes 3*7, which makes a mixed dozenal and decimal world harder than it needs to be. But this raises the interesting question of what a logical dozenal set of glyphs might look like.
I've amused myself with thinking about this, and I have a proposal for a set of glyphs that are (mostly) unused, easy to learn, quick to draw, and are much more logical than existing ones. All have a vertical line in the middle. At the top there is a choice of a hook starting on the left, a straight end, or a hook starting on the right. At the bottom there is a choice of a hook ending on the left, straight down, on the right, or bending right to cut across the vertical line. This gives 12 possibilities. By incrementing the bottom first, and the top when you get a carry, you get a sort of a 3,4 base. Which means that a glance at the glyph tells you immediately its sign mod 4. A glance at the top tells you whether it falls in the range 0-3, 4-7 or 8-11.
So instead of the current system of 10 random symbols to memorize, you get two simple rules. While at first it seems bizarre, it grows on you quickly. And it gives you a very quick way to tell whether a given number is written in decimal or dozenal. To get a sense what it looks like, take a look at the 12 times table (my apologies for the handwriting - my none too good penmanship gets worse when I'm using a mouse pad to draw with):

Note in particular how regular the patterns are for 2, 3, 4, and 6. Doesn't this look easier to memorize than the decimal times table? As an exercise try writing out your favorite sequences. Whether you're writing out squares, powers of 2, or primes you'll see that more patterns leap out at you in dozenal, making them easier to learn.
So there is my humble contribution to a dozenal future.
(In other news, I now have a patent to my name, though in fact it is owned by a previous employer. By the standards of the patent system, it is not a particularly bad patent. But if I had my druthers, it would have never been filed. Ah well.)
When the French created the metric system, they based it on factors of 10 everywhere. They even went so far as to try to measure angles in gradians (a quarter circle had 100), and to use decimal time. The world as a whole rejected decimal angles and times, but has adopted decimal metric everywhere else. Which is very convenient for scientists, but is hard to divide into thirds and somewhat inconvenient for quarters. Furthermore the persistence of both systems results in occasional annoyances like the fact that in daily life we usually prefer to measure speed in km/h, but for energy and power calculations the units only work out properly if you measure in m/s. Resulting in an annoying factor of 3.6 that shows up converting between them.
The other day I read An Argument for Dozenalism that made many of these arguments. Nothing new. However it made the interesting point that ideally a dozenal arithmetic would have its own set of glyphs. It suggested that 0 and 1 could be kept, but everything else should be changed. In my opinion that argument is correct, it is very confusing if 21 sometimes is 5*5 and sometimes 3*7, which makes a mixed dozenal and decimal world harder than it needs to be. But this raises the interesting question of what a logical dozenal set of glyphs might look like.
I've amused myself with thinking about this, and I have a proposal for a set of glyphs that are (mostly) unused, easy to learn, quick to draw, and are much more logical than existing ones. All have a vertical line in the middle. At the top there is a choice of a hook starting on the left, a straight end, or a hook starting on the right. At the bottom there is a choice of a hook ending on the left, straight down, on the right, or bending right to cut across the vertical line. This gives 12 possibilities. By incrementing the bottom first, and the top when you get a carry, you get a sort of a 3,4 base. Which means that a glance at the glyph tells you immediately its sign mod 4. A glance at the top tells you whether it falls in the range 0-3, 4-7 or 8-11.
So instead of the current system of 10 random symbols to memorize, you get two simple rules. While at first it seems bizarre, it grows on you quickly. And it gives you a very quick way to tell whether a given number is written in decimal or dozenal. To get a sense what it looks like, take a look at the 12 times table (my apologies for the handwriting - my none too good penmanship gets worse when I'm using a mouse pad to draw with):

Note in particular how regular the patterns are for 2, 3, 4, and 6. Doesn't this look easier to memorize than the decimal times table? As an exercise try writing out your favorite sequences. Whether you're writing out squares, powers of 2, or primes you'll see that more patterns leap out at you in dozenal, making them easier to learn.
So there is my humble contribution to a dozenal future.
(In other news, I now have a patent to my name, though in fact it is owned by a previous employer. By the standards of the patent system, it is not a particularly bad patent. But if I had my druthers, it would have never been filed. Ah well.)
Subscribe to:
Posts (Atom)