Showing posts with label model. Show all posts
Showing posts with label model. Show all posts

Rowing Across an Ocean of Contours

3 comments:
Today I'm very happy to feature a guest post from my friend Angie Pendergrass, a PhD student in the Atmospheric Sciences department here at UW. She's got the mother of all side-projects: using her data analysis and climate modeling skills to help predict ocean currents for a group of crazy brave rowers...




Imagine you're rowing a boat across the ocean. You can't row very fast, just 3 knots if you try your hardest (that's a little slower than walking speed.) You have 4000 nautical miles nautical miles of sea to cover. Scattered across the ocean surface, you know there are some favorable currents and some eddies, spinning whirlpools of water about 50 nautical miles across, with current up to at least one knot. They can help you if you can ride them favorably, or they can suck -- a lot -- if they're against you. You really need to know what these eddies are doing!

Luckily you have a team of weather forecasters and a navigator ashore helping you out. There are actually some guys doing this row, called OAR Northwest. They set out from Dakar, Senegal in late January for Miami, Florida, and now they're halfway across the middle of the Atlantic; check them out at oarnorthwest.com. Your guest-author is their lead weather forecaster.

So now you should be super excited to hear that the National Center for Environmental Prediction (part of NOAA) runs a model of the ocean that diagnoses the ocean surface currents. Great, this is perfect! Let's take a look!

Two Distorted Stars

3 comments:

I have spent the last week in Alaska at the 220th AAS conference. It was a blast, but did not lend itself to blogging...

By way of an apology, here is a (gratuitous) animation of a binary star system orbiting that I created for my research using the PHOEBE software package. The animation was created by stitching 10ish frames (generated in IDL) together in gifsicle, which is kind of an old-school way of doing this. This visualizes two stars orbiting each other in a "semi-detached" configuration.


The smaller star is about 70% the mass of the bigger one. They are both distorted, with a characteristic tear-dropped shape, due to the tidal gravitational force between them. We will be discussing many such systems at the forth-coming "Al-Fest" workshop at UW next month.

This was a simple test case animation I made, and is not the final configuration for the short period binary I recently presented at the conference. (The paper should be submitted this week has been submitted!)

Nothing else too fancy to say about it. I think it's quite mesmerizing, and utterly fascinating to consider that two stars can be orbiting each other so quickly and closely. In the case of my short period object, composed of two M dwarfs, they go through a full orbit once every ~4 hours!

Like with my previous Plots as Art post(s), sometimes the research visualizations are just too fun not to share.

Never Take SR-520

8 comments:
Be sure to subscribe for updates on this and all my other data analysis projects!


Fuel prices have always been higher in Seattle than in many other places in the country. I've been paying around $4.30, and looking at this map from AAA that seems about average for the west coast.
AAA Fuel Gauge Report Map

This year the State Route 520 floating bridge in Seattle become a toll road, after several delays. Besides enabling the automatic tolling service 8 months behind schedule, it's been a rough year for 520. The tolls have been occasionally overcharging drivers, the bridge has been closed many weekends for maintenance (see picture below), and a yacht got wedged under the highway. Yikes.

SR520 closed for weekend
With gas prices climbing, and the 520 tolls due to increase by 2.5% starting July 1st, I thought it high time to finish a post I'd started working on months ago! The premise is simple:
Is it cheaper to take 520, or drive around on I-90?

Tolling prices for weekdays (black) and weekends (blue)

To address this question, we must first understand what the tolling prices are like. Above I have reproduced the weekday and weekend prices as a function of time. You can see the variable price scheme, whereby the drivers in morning and evening rush-hours are relatively fleeced. On the weekends it's a single-peaked distribution, cashing in on (for example) Saturday shoppers headed from Seattle to the popular Bellevue Square shopping mall.

This flat cost is the first part of the equation. Next we need to estimate how fuel efficient your car is, so we can calculate the price of driving the different length routes.




Here I'm showing the fuel efficiency of a car as a function of its speed. The black line describes this efficiency for a fairly ideal passenger car, like my 1999 Honda. The blue line has been reduced in efficiency by 20%, and I believe is more representative of a typical "light SUV", such as the extremely popular (in the northwest) Subaru Outback. I'll be assuming the SUV case below. I don't really know how accurate this data is, but I've assembled it from fueleconomy.gov and I'll let you decide if you think this is a reliable or impartial source.

The last piece of the puzzle is a route to consider! For this example I'll be traveling from UW to Bellevue Square mall. Route A is the direct across 520, route B is the long drive around I-90, via I-5 and I-405. I played with other routes as well, but this is the simplest and most dramatic.

Route A: the direct trip
Route B: the long way around
a


So the game works like this: At every hour of the day, for every speed between 0 and 80mph, I have calculated the cost of going to Bellevue via both routes. The reason I have gone through the silly process of computing cost for every road speed is that I was curious if the distance would be the deciding factor at certain speeds. As such, I apologize if the result figure is complicated. 

Weekday price difference grid, assuming $4.20/gal fuel. Black lines highlight the price
difference over time at 60mph, the nominal speed limit on both roads.
You can almost read off the data from the grid. The color indicates the difference in price between driving around and taking 520. Red indicates that 520 is more expensive. Green to black is I-90 being more expensive, 520 cheaper. The color-bar (legend) at top was capped at $1, but the maximum price difference was actually around $1.80.

The result gives rise to the pointed name I chose for this post: you should (almost) never take SR 520 if you want to save money. The exceptions are:
  • late at night, when the toll is not in effect
  • if you're driving 5mph
  • if your time is worth more than the amount you save driving around. As a graduate student, the choice is laughably clear.
During the off-hours (11pm-5am) you can see the shape of the fuel economy curve quite nicely. Some of that effect can also be seen at the high-speed end. A reminder to you all: going 55mph saves lives and gas, which in turn saves other lives.


It may occur to you (or, it did to me at least) that this result depends on what the price of gas is, and you'd be absolutely correct. So I turned that knob in the model, and found that at ~$7/gallon gas it becomes cheaper to take 520 at all times except rush hour.

The result also depends on your fuel economy in as similar way. I've only explored the case of a small SUV, but consider a semi truck that gets 6-8 miles per gallon! In almost every conceivable scenario it would be cheaper to save a gallon or two of gas and take 520.

Maybe I'm an optimist, or a fool, but it seems like a reasonable proposal for the toll prices would have been to make them price-neutral to driving around. Tolling also increases traffic on other roads, decreasing the profitability some. I did enjoy seeing that the tolls don't affect low income families greatly. I read through a few of the initial reports released about the 520 tolls, but I'm very interested in looking at the traffic statistics and financial data to see if it meets predictions.

Overall I'm struck with how different the data analysis seems to be in this field. Figures don't seem precise, but rather very bold and contrasting. Sometimes this is effective and accurately tells a story better than a very precise figure. Frequently I find the plots burdensome, but that's the subject for another post...

Coffee: Two Minutes from Anywhere

2 comments:
Be sure to subscribe for updates on this and all my other data analysis projects!

I live in Seattle, the self proclaimed "coffee capitol" of the USA. We have a Starbucks on every corner downtown, independent cafes all over, and even some Starbucks incognito as indie coffee stores. It's a wonderfully caffeinated culture we're brewing out here.

A long standing joke among my friends at the University of Washington (UW): we drink so much coffee that we put a cafe in nearly every building on campus. This isn't quite true, of course, but we do have many!

While driving home from school today it occurred to me to consider the joke in a different direction: how far on UW's campus can you get from a cafe? Are you ever more than 2 minutes away from a coffee stand?

So I set about finding the answer!

Water Water

1 comment:
I found a link while browsing reddit this afternoon (from r/dataisbeautiful) that pointed to a community data visualization challenge.  The source for the data was a "Global Water Experiment", which provided some basic measurements of water characteristics, such as pH.

Though the data challenge had passed its deadline by a couple weeks, I was still intrigued, and so I used a couple free hours while some code was running on my work machine to play with this "Global Experiment".

Aside: It is becoming clear to me that I need to learn some new data visualization software, and I haven't been very impressed with most graphics I've seen from python. R sounds like a cool option, it's free and widely used... we'll see what I get in to.


I downloaded the .xls file, and while looking through the data on pH levels a question popped in to my head!


Question: does the "financial prowess" of a country correlate with the quality of its drinking water as tracked by pH levels?


This really reminded me of the (now famous) TED talk by Hans Roling, and his cute data visualization that everyone gushed over a few years ago - and it really is a cool talk, btw.  Some quick searching online suggests that this sort of effect has been investigated before, but results have been largely inconclusive, or consistent with no correlation.

The challenge data provided me with over 2000 pH measurements from almost a hundred countries, though the sampling seemed sporadic between countries. I had to spend some time "cleaning" the data file by making names of countries uniform (e.g. USA became United States) and getting rid of non-standard characters.

The financial numbers came from Gross Domestic Product (GDP) data, provided courtesy of the CIA as it so happened!

So armed with my two sources of data, I set about the ever-fun task of string matching. It's somewhat of a clunky operation in IDL, and I always recommend people try a few things when doing this:

  1. trim any leading/trailing spaces using STRTRIM(inputstring,2)
  2. convert everything to lower case using STRLOWCASE(inputstring)
  3. if things vary too much, you can try matching over only part of the string using the vector functionality of STRMID()
I matched the GDP and pH data, and then selected only the fresh water samples. I was a bit hasty with the matching, and probably threw out by accident some of the pH sampling. I also missed some country name matching, no doubt.

For each country with GDP and pH data, I measured the mean (average) and standard deviation of the pH samples, with Std Dev only calculated for countries with 2 or more samples. The immediate red flag that should go up in your mind (or certainly did for me) was: don't some countries have drastically different water environments that are being tested? Absolutely! As I mentioned above, this study didn't seem to guarantee any certain degree of accuracy or completeness. Furthermore, as a good friend/mentor of mine once reminded me when I was an undergrad, science requires error bars! These data have none that I could find.

The Std Dev may or may not be terribly useful, but I felt the means were quite illuminating. Here is the figure:

Figure 1: Top) Mean temperature reported vs GDP. Bottom) Mean water pH level reported, with standard deviations for each country shown as error bars. A linear least squares fit is provided in red.
As you can see from this figure there is a weak trend with GDP. Also curious to me was the apparent anti-correlation between temperature and GDP. Evidently, if you want to live in a well-to-do country, live somewhere cooler!

That furthest-right data point: the good 'ol USA of course! Several other name-brand western countries comprise the right-most portion of this figure (such as the UK, Canada, Australia). Noticeably absent from this data set: freshwater measurements from China.

Conclusion: There seems to be a weak trend, with water pH levels for more "developing" countries preferentially basic. Intermediate economies seem to be quite spread in pH levels. The most powerful nations continue the rough trend seen over four orders of magnitude in GDP: a decline from alkaline towards pure water.

The rough trend may seem encouraging for politicians, but the significant scatter represents both the intrinsic noise in the data, and a wider ecological issue. Consider that the USA has the highest economy listed by a wide margin, as well as one of the widest standard deviations in pH. Several countries with GDP's smaller than most American states have markedly better water quality. This supports the indication (mentioned in the abstract linked near the beginning) that it is not the GDP which contributes most to water quality. I am left thinking of a few better candidates: social/political forces, environmental regulations, natural resources, or maybe just number of trees...

That's enough wild (data-less) speculation for tonight I'd say.