ZIP Code Density

No comments:
Color/darkness represents the number of ZIP codes within each pixel.

I've been playing with some neat data from the 2010 Census, and in honor of Memorial Day I thought this was a super fun plot to show. The actual post that I'm working on related to this will probably take a long time to complete, as it requires I gather a ton of data...

 I'm using the "CubeHelix" color scheme (shameless promotion: I wrote the first version for IDL). This scheme is great because it de-saturates from color to black/white while correctly preserving brightness. It can be rainbow colors too, not just monotone like above.

For example, here is the same figure, but with the color turned off (using Apple's Preview app). Great for making figures that look good both online (where color printing is free) and in black/white print (e.g. ApJ charges $350 per color figure, last I checked)



Never Take SR-520

8 comments:
Be sure to subscribe for updates on this and all my other data analysis projects!


Fuel prices have always been higher in Seattle than in many other places in the country. I've been paying around $4.30, and looking at this map from AAA that seems about average for the west coast.
AAA Fuel Gauge Report Map

This year the State Route 520 floating bridge in Seattle become a toll road, after several delays. Besides enabling the automatic tolling service 8 months behind schedule, it's been a rough year for 520. The tolls have been occasionally overcharging drivers, the bridge has been closed many weekends for maintenance (see picture below), and a yacht got wedged under the highway. Yikes.

SR520 closed for weekend
With gas prices climbing, and the 520 tolls due to increase by 2.5% starting July 1st, I thought it high time to finish a post I'd started working on months ago! The premise is simple:
Is it cheaper to take 520, or drive around on I-90?

Tolling prices for weekdays (black) and weekends (blue)

To address this question, we must first understand what the tolling prices are like. Above I have reproduced the weekday and weekend prices as a function of time. You can see the variable price scheme, whereby the drivers in morning and evening rush-hours are relatively fleeced. On the weekends it's a single-peaked distribution, cashing in on (for example) Saturday shoppers headed from Seattle to the popular Bellevue Square shopping mall.

This flat cost is the first part of the equation. Next we need to estimate how fuel efficient your car is, so we can calculate the price of driving the different length routes.




Here I'm showing the fuel efficiency of a car as a function of its speed. The black line describes this efficiency for a fairly ideal passenger car, like my 1999 Honda. The blue line has been reduced in efficiency by 20%, and I believe is more representative of a typical "light SUV", such as the extremely popular (in the northwest) Subaru Outback. I'll be assuming the SUV case below. I don't really know how accurate this data is, but I've assembled it from fueleconomy.gov and I'll let you decide if you think this is a reliable or impartial source.

The last piece of the puzzle is a route to consider! For this example I'll be traveling from UW to Bellevue Square mall. Route A is the direct across 520, route B is the long drive around I-90, via I-5 and I-405. I played with other routes as well, but this is the simplest and most dramatic.

Route A: the direct trip
Route B: the long way around
a


So the game works like this: At every hour of the day, for every speed between 0 and 80mph, I have calculated the cost of going to Bellevue via both routes. The reason I have gone through the silly process of computing cost for every road speed is that I was curious if the distance would be the deciding factor at certain speeds. As such, I apologize if the result figure is complicated. 

Weekday price difference grid, assuming $4.20/gal fuel. Black lines highlight the price
difference over time at 60mph, the nominal speed limit on both roads.
You can almost read off the data from the grid. The color indicates the difference in price between driving around and taking 520. Red indicates that 520 is more expensive. Green to black is I-90 being more expensive, 520 cheaper. The color-bar (legend) at top was capped at $1, but the maximum price difference was actually around $1.80.

The result gives rise to the pointed name I chose for this post: you should (almost) never take SR 520 if you want to save money. The exceptions are:
  • late at night, when the toll is not in effect
  • if you're driving 5mph
  • if your time is worth more than the amount you save driving around. As a graduate student, the choice is laughably clear.
During the off-hours (11pm-5am) you can see the shape of the fuel economy curve quite nicely. Some of that effect can also be seen at the high-speed end. A reminder to you all: going 55mph saves lives and gas, which in turn saves other lives.


It may occur to you (or, it did to me at least) that this result depends on what the price of gas is, and you'd be absolutely correct. So I turned that knob in the model, and found that at ~$7/gallon gas it becomes cheaper to take 520 at all times except rush hour.

The result also depends on your fuel economy in as similar way. I've only explored the case of a small SUV, but consider a semi truck that gets 6-8 miles per gallon! In almost every conceivable scenario it would be cheaper to save a gallon or two of gas and take 520.

Maybe I'm an optimist, or a fool, but it seems like a reasonable proposal for the toll prices would have been to make them price-neutral to driving around. Tolling also increases traffic on other roads, decreasing the profitability some. I did enjoy seeing that the tolls don't affect low income families greatly. I read through a few of the initial reports released about the 520 tolls, but I'm very interested in looking at the traffic statistics and financial data to see if it meets predictions.

Overall I'm struck with how different the data analysis seems to be in this field. Figures don't seem precise, but rather very bold and contrasting. Sometimes this is effective and accurately tells a story better than a very precise figure. Frequently I find the plots burdensome, but that's the subject for another post...

Rumor Names - 2012

No comments:
A few months ago I put up a word cloud that analyzed the most common name for people getting postdoc jobs in astronomy in 2011 (according to the Rumor Mill). The result: if you wanted to get a job in astronomy last year, you had better be named David...

The 2012 job season has all be ended for the Academe. I'm happy to say it has been kind to many of my friends. Despite the foreboding messages from some talking heads, I know of none with PhDs who are requiring food stamps.

So I went to the rumor mill once more and grabbed the names of the employed elite. This year I actually tried to clean the wordle of the frequent garbage that would get included, but a few errant words slipped through.
Frequency of names from the Postdoc Jobs Rumor Mill, as of  2012 May 21
The winner this year was "Yan", which is actually due to both Tze Yan Lam and Francis-Yan Cyr-Racine being offered multiple positions each. Even more interestingly both Yan's work on cosmology theory, and were offered some of the same positions. I don't believe they have collaborated, but it seems fitting that they should.

Here's a breakdown of the contenders this year...

  Yan       11
  Laura     9
  David     8
  Michael   8
  Fumagalli 7
  Michele   7
  Colin     7
  Hamaus    7
  Wenjuan   7
  Voort     7
  Freeke    7
  Fang      7

That was all the results up to 7 hits. There was a big pile up at around 4-6 hits, which is why the word cloud looks so homogenous in word size. Odds are still very good that if you're named David you can find work in Astronomy! The winning first name this year, however, was Laura.

Fun stuff!

Email History

No comments:
I came across this neat article in WIRED a couple months ago by Stephen Wolfram (the man who brought you Mathematica, WolframAlpha...). In it, Dr Wolfram discusses some of the impressive and mildly obsessive amounts of data about himself that he has kept for over 10 years. It's a fascinating look at how we truly can be statistical beings.

He naturally makes some dandy plots to showcase this information. The first figure in the article, showing the date vs time of day for every email since 1990, really caught my eye and I've wanted to make it ever since.

Alas, I did not have the resources available to the good Dr Wolfram, and have not been able to keep my entire academic email history. GMail promised such a revolution, to "Never Delete Emails", and I have every message sent through that service since late 2005 still online. I am in the process of downloading them all. Until my GMail finishes downloading, I decided to just make the figure with what I had available on my laptop: my academic email for the last year.

This is my recreation of that super-neat figure, showing the time of day that I send emails. Even with my limited number of emails, this is a really cool figure! Most of the early AM email spurts are from days when I'm observing. I otherwise appear to be a regular creature of habit, at least with my work email.  Let's dive in just a little further...

Here I very simply show the histogram of how many emails I sent each day of the week. I try to take Saturdays off, and you can literally see the burn-out happening Thursday-Friday.

Here we have the normalized cumulative distribution of emails, showing the running total of emails normalized to 1. Besides the artifact on the first day in 2011, when I first set up the new UW email server on my laptop, the rate of email has been surprisingly constant/linear!

Lastly, here is a histogram of when I send email throughout the day. This highlights my daily routine: Sleep around 1 AM, up at 7-8, lunch around 12:30, dinner at around 7pm (19:00).

My takeaway message: I have been doing a good job of keeping a regular schedule for the past year, something I struggled with at times in college, and is often a problem for graduate students. This last year has been the first time since I was 6 years old that I have not taken classes, and I was justifiably concerned with my ability to keep a regular schedule of waking up early and working. These results are encouraging... now if only I could plot how much actual work I've been able to do as a function of time. That might be a less enjoyable result. In the finest academic tradition, I leave that as an exercise for the reader.