Showing posts with label observations. Show all posts
Showing posts with label observations. Show all posts

Alaska LEOS and rare Mourning Doves

I finally participated in my first LEO Webinar and had a great time.  I'll be calling in for more of them as they come up monthly.

LEO is the Local Environmental Observers program/project in Alaska.  The principle being, the people actually living in an area are the prime observers for what is going on.  This includes keeping an eye on birds, among many other things.  The title comes from the September 13, 2013 observation by Richard Kuzuguk of a Mourning Dove in Shishmaref, AK.  Very unusual up there.  I don't have that species down here, but in general, they're very common here.  (Common as in 'wake up light sleepers'.)

For an idea of the rarity of Mourning Doves in Western Alaska, take a look at the distribution map at Wikipedia.

Starting a bestiary of oscillations and cycles

A bestiary originally was originally a book of pictures and descriptions, often with morals attached, of animals.  Well, that was the middle ages.  The version I've got in mind is one describing the more or less regular oscillations or cycles in the earth system, its orbit, and the sun.  For now, I'll describe just the period and its name and invite you to add to the list.  Also, I won't worry about whether the named thing is a proper oscillation (such as tides) or more of an index that may not have any particular period (PNA).  The later rendition will have some discussion of what happens in each and concern about whether the variation is a real thing or just an artefact of how people looked at the data.  For those who'd like to jump straight to discussion of weather cycles directly, I'll suggest William Burroughs' Weather Cycles, Real or Imaginary

As always, you're encouraged to add your own suggestions!
Around a day

12h 25 min (12:25) -- Lunar semidiurnal tide
23:56 -- Sidereal day
24:00 -- Mean Solar day (1 dy)
24:50 -- Lunar diurnal tide

Around a month
13.66 dy -- Lunar fortnightly tide
27.32 dy -- Lunar monthly tide
40-60 dy -- Madden-Julian Oscillation (MJO)

Around a year
365.2422 dy --  Tropical Year (equinox to equinox) (1 year for later)
365.2564 dy -- Sidereal year [see comment]
365.259 dy -- Anomalistic Year (perihelion to perihelion)
~433 dy -- Chandler Wobble

A few years
~26 months -- Quasi-biennial Oscillation (QBO)
2-7 years -- El-Nino/Southern Oscillation (ENSO), Southern Oscillation Index (SOI)
~4 years  -- Antarctic Circumpolar Wave
~3.75 years -- Rossby-Kelvin wave in North Pacific [see comment]
8.85 years -- Lunar perigee
18.6 years -- Precession of the Lunar Node

Many years
(See http://http://www.arctic.noaa.gov/essay_bond.html for some discussion)
--- -- Arctic Oscillation (AO)
--- -- Antarctic Oscillation (AAO)
--- -- North Atlantic Oscillation (NAO)
--- -- Pacific/North America Pattern (PNA)
20-30 years -- Pacific Decadal Oscillation (PDO)

~11 years -- Sunspot cycle
~22 years -- Solar cycle (each sunspot cycle is opposite magnetic polarity)
~88 years -- Gleissberg cycle (clumping of solar cycles)


Long period
19-23,000 years -- Milankovitch Cycle -- Precession of the equinoxes
~41,000 years --Milankovitch Cycle -- Tilt of the earth
~100,000 years -- Milankovitch Cycle -- Eccentricity of the earth's orbit
~400,000 years -- Milankovitch Cycle -- Eccentricity of the earth's orbit

Very long period
30 million years --  Oscillation of solar system above/below the plane of the galaxy
230 million years -- Solar system orbit of the galaxy [see comment]
400 Million years -- Supercontinent cycle

Constructing an analysis 1: Drop in a bucket

'Analysis' is what we call an attempt to represent the state of the atmosphere/ocean/sea ice/... given a set of observations. One such analysis is the global surface air temperature analysis. That, then, spawns efforts to find a global mean temperature, or global mean temperature trends, and so forth. Several of the recently-added blogs aim to study that, in one way or another. That particular one is not my interest in two different ways.

One is, I'm an oceanographer, so I'm more interested in a sea surface temperature (sst) analysis. The other is, most of the interest in the surface air temperature analysis seems to come from its role as a detector of climate change. On the scale of things, I consider this the second weakest climate change indicator. The only thing weaker, in my view, is the so-called 'Hockey Stick'. But enough raw opinion.

Regardless of what it is you're trying to analyze, and what your reason for doing so is, there are quite a few ways of setting about doing so objectively. The fact that there are many makes this the first of something like eight notes I'll be writing up on the idea. There turn out to be many different ways of making an analysis, each objective, each with strengths, each with weaknesses.

The simplest one, if not as simple as you might think, is the 'drop in a bucket' method.
First, a bit of language.  We typically divide the earth's surface in to a bunch of boxes/cells.  Also typically, they're some number (or fraction) of degree latitude by so many degrees (or fraction) of longitude.  Depending on which sst analysis I'm looking at, a cell can be anything from 5 degrees on a side to 1/100th of a degree on a side.

The basic idea for drop in a bucket is very simple -- if you have temperature observations in a cell, you use them to find the temperature of (analyze) that cell.  If you have no temperatures in a cell, then you have no analysis for that cell.  So 0 observations in a cell is very easy -- you report no analysis.  1 observation is also very easy, your analysis temperature is the temperature from that one observation.

But what about having more than 1 observation in a cell?  The very simplest thing to do is just average all of them.  We just blindly treat all observations as being equally good.  Hmm.  That sounds a bit problematic.  Some observing methods are better than others, after all.  The quality of the observations is described by the standard error, which is the standard deviation between the true value and the observed value -- computed after you have many such observation to truth comparisons.

For typical drifting buoys and satellite methods, this is about 0.5 degrees.  For ships, let's say 1 degree.  This being science, of course I mean degrees C.  One could pursue this to substantial complexity, as it's probably the case that every type of buoy has a somewhat different standard error, different ship observing methods have different standard error, and the different satellites and satellite methods have still other standard errors.  Life is probably no simpler for surface air temperature observations; and for rain it's even harder.

For the sake of illustration, let's consider a cell with a buoy that observed a temperature of 25 C, and a ship in the same area that observe 26 C.  In blind averaging, we'd treat them as equal, and give our analysis as 25.5 C.  But ... the buoy is a better observer than the ship.  Shouldn't our analysis be closer to the buoy?  Maybe we should just throw out the worse observer?

Probably not.  The observations have a distribution of likelihood (which is not the vertical axis! beware!) around the value they report.  For each observation, the most likely value is what is reported.  But those standard errors give us a curve of likelihood.  The ship is in orange, the buoy in blue, and a third thing in black:
The ship (orange) has the large standard error, so it's a very wide curve.  It's odd, but true, that the ship observation is more in agreement with a claim of 23 than the buoy observation, even though the buoy observation is colder.  The thing is, the buoy observation is less uncertain.

The third curve is where we get to the creative part.  What it is, is that I've multiplied the likelihood curves for the buoy and the ship.  The result is a joint likelihood.  The peak of the curve is our point of maximum likelihood.  It's a temperature of 25.2 C.  You can, in principle, do this sort of thing graphically regardless of how many observations you have.  It gets tedious and ugly, of course.  That's why we invented mathematics.  In this case, calculus.  (See bottom for the gory math details). 

We also see that our resulting estimate, the maximum likelihood curve, is narrower than the two original ones.  The more observations we have, the better our resultant estimate -- even better than the original observations.  This is the same sort of thing we saw result in How can annual average temperatures be so precise?

Our estimate based on considering the quality of the different observing platforms is 25.2, rather than the 25.5 of treating them as equally good.  When we're looking to deal with climate change and detecting small differences over time, it's obviously important to pay attention to this sort of change.  If you change from one to the other, there's a 0.3 C change -- not because of climate, but because you changed your methods for filling cells.    (Not a mistake I think anyone has made, but a heads up if you are getting started.)

The general term for this sort of thing (treating some observations as better than others) is 'weighting'.  We give more weight to some sources than others.  I'll give the exact method at bottom for the mathy folks.  Different methods that we'll be getting to will do their weighting in different ways.

Now let's go back to thinking about the general approach.  We select boxes of some size, and then if there's an observation in the box, then we say that we know nothing about what's going on in the box.  There are about 5000 drifting buoy observations per day.  If the cells are 5 degrees on a side, and the buoys are distributed randomly, there are about 3 observations per grid cell and we are doing pretty well.  On the other hand, if our cells are 1/100 degree on a side, then probably only one cell in 80,000 has an observation.  How did we go from knowing most of the globe pretty well (3 observations, obs, per cell on average) to knowing almost nothing?  On the other hand, for every cell you say you know something, you definitely have at least one observation in support, and you haven't made any assumptions (at least not past selecting the cell size).

But suppose you are running a numerical weather prediction (NWP) model.  You can't accept an 'I don't know' for starting your prediction.  Consequently the people involved in NWP were early people to develop more advanced methods.  That's next.

Methods (so far):
1a) Drop in bucket, blind averaging
1b) Drop in bucket, maximum likelihood averaging


Gory math details:

To find the maximum likelihood estimate for temperatures given N observations,  multiply together your N curves, each of the form exp( - (T-T_i)^2/s_i^2), and find the T for which this is a maximum.  T_i is the ith observation, s_i is its standard error.

The maximum likelihood value is sum(T_i / s_i^2) / sum(1 / s_i^2).

I'll leave it for those interested to compute the standard error of the maximum likelihood estimate itself.

Were the 70s cold?

I was surprised to see that the 1970s weren't particularly cold.  My surprise is partly because where I lived (Chicago area) we were busy setting all-time records for cold, and that was true for much of the US and across to the UK. 

The other part of the surprise is that it's common to hear people (see them write) something on the lines of "Of course we're seeing a warming since the 70s; it was cold in the 70s!"  Surely someone along the way did their homework and checked out what the global temperatures were?

Fortunately, if we're looking at science, we don't have to assume that other people did their work, or did it correctly.  The alternate word for it is, skepticism.  Real skeptics don't make those assumptions, they do the work themselves.  The fact that it's work also explains why there are a lot of fake skeptics -- it's much easier to pick the answer you like and reject everything else.

So let's apply some real skepticism and ask what was really going on with temperatures in the 1970s.
I'm working from the NCDC global mean temperatures.  I'll take the average of each decade they give, and plot that:

That's rather disappointing.  The 1970s were the 4th warmest (of 12) decades in that record.  Actually a little worse, as you'll notice I don't have the decade starting in 2000 computed.  (I downloaded the file in 2009, it's gotten warmer since then). 4th warmest of 13 decades, behind only 1940s, 1980s, 1990s, and 2000s. But -- be a proper skeptic -- compute for yourself the average for January 2000 through December 2009.  I got +0.35 for the 1990s.

My disappointment is twofold.  First off is that the decade that I remember as being so very cold, wasn't.  In my area it was, and even for a good distance away.  But the entire US is only about 2% of the globe.  The global average could easily be quite different, and turns out that it was.  The 1970s were a relatively warm decade.

The second part of my disappointment is that those people calling themselves skeptics are clearly not skeptics.  They never computed decadal temperatures, and never looked to see whether the 1970s were particularly cold.  Not only were they not cold, but they were warmer than the two decades before them.  Real skeptics would not claim that they were cold in talking about climate change.

Data set reproducibility

Data are messy, and all data have problems.  There's no two ways about that.  Any time you set about working seriously with data (as opposed to knocking off some fairly trivial blog comment), you have to sit down to wrestle with that fact.  I've been reminded of that from several different directions recently.  Most recent is Steve Easterbrook's note on Open Climate Science.  I will cannibalize some of my comment from there, and add things more for the local audience.

One of the concerns in Steve's note is 'openness'.  It's an important concern and related to what I'll take up here, but I'll actually shift emphasis a little.  Namely, suppose you are a scientist trying to do good work with the data you have.  I'll use data for sea ice concentration analysis for illustration because I do so at work, and am very familiar with its pitfalls.

There are very well-known methods for turning a certain type of observation (passive microwaves) in to a sea ice concentration.  So we're done, right?  All you have to do is specify what method you used?  Er, no.  And thence comes the difficulties, issues, and concerns about reproducing results.  The important thing here, and my shift of emphasis, is that it's about scientists trying to reproduce their own results (or me trying to reproduce my own).  That's an important point in its own right -- how much confidence can you have if you can't reproduce your own results, using your own data, and your own scripts+program, on your own computer?  Clearly a good starting point for doing reliable, reproducible, science.



This turns out, irrespective of any arguments about the honesty of scientists, to be a point of some challenge even as software engineering.  Some of this was prompted by a discussion I had some months ago at a data-oriented meeting -- where someone was asserting that once the data were archived, future researchers could 'certainly' reproduce the results you got today.  I was not, shall we say, impressed.

We'll start by assuming that the original data have been archived.  (Which I've done for my work, at least for the most recent period.)  Ah, but did you also archive the data decoder?  Turns out that even though the data were archived exactly as originally used, the decoder itself is allowed to give slightly different results when acting on the same data (or maybe there was a bug in the old decoder that was fixed in the newer one?).  So, even with the same data, merely bringing it out of the archive format in to some format you can work with can introduce some changes.  Now, do you archive all the data decoding programs along with the data?  Use the modern decoders? (but when is modern?  If this year's decoder gives different answers than 5 years ago, or than 5 years from now, what should the answer be today?)

Having decoded the data to something usable, now we run our program that translates the data we have in to something that is meaningful.  In my case, this means translating 'brightness temperatures' (themselves the result of processing the actual satellite observations in to something that is meaningful to people like me) in to sea ice concentrations.  The methods are 'published'.  Well, some of the method is published.  The thing is, the basic algorithm (rules) for translation are published -- it's the NASA Team algorithm from 1995 through 23 August 2004, then my variation from then to August 2006, and then my variation on top of NASA Team2 from there to the present.  One issue being, my variation is too minor to be worth its own peer-reviewed scientific literature (though I confess I've been reconsidering that statement lately, as more trivial-to-me papers are being published).  So, where, exactly is the description of the methods?  Er.  In the programs themselves.

That's not a problem in itself.  I have, I think, saved all versions of my programs.  Related, though, is that the algorithms don't really stand on their own.  There is also the matter -- seldom publishable in the peer-reviewed literature, but vital to being able to reproduce the results -- of what quality control criteria were used, and what, exactly, the weather filtering was.  To elaborate: As I said, data are messy and ugly.  One of the problems is that the satellite can report impossible values, or at least values that can't possibly correspond to an observation of the sea ice pack.  Maybe it's a correct observation, but of deep space instead of the surface of the earth.  Maybe the brightness temperatures are correct, but the latitude and longitude of the observation are impossible (latitude greater than 90 N, for instance).  And maybe just there was a corruption in the data recorder and garbage came through.  In any case, I have some filters to reject observations on these sorts of grounds before bothering the sea ice concentration algorithm with them.  And then there's the matter of filtering out things that might be sea ice cover (at least the algorithm thinks so) but which are probably just a strong storm that's got high rain rates, or is kicking up high waves ('weather').

But, of course, these quality control criteria, and weather filter, have changed over the years.  Again, not publishable results on their own, but something you need in order to reproduce my results.  And, again, the documentation, ultimately, is the program itself.  Since I've saved (I think) all the versions of the relevant program(s), you might figure we're home free -- fully reproducible results.

Or, at least the results would be exactly reproducible if you also had several other things.

One of them is, the programs rely on some auxiliary data sets.  For instance, the final results like to know where the land is in the world.  So I have land masks.  If you want my exact results, you need the exact set of land masks I used that day.  Again, I've saved those files.  Or at least I think I have.  As many a person has discovered the hard way at home -- sometimes your system back up won't restore.  Or what you actually saved wasn't what you meant to save.

It's worse than that, though.  You probably can't run my programs on your computer.  At least not exactly my programs.  I wrote them in some high level language (Fortran/C/C++) and then a compiler translated my writing in to something my computer (at that time) can use.  The exact computer that I was running on that day.  Further, there are mathematical libraries that my program uses (things to translate what, exactly, it means to compute the sine of an angle, or a cosine), that were used the day the program originally ran.

What you can do is compile my program's source code (probably, I'm pretty aggressive about my programs being able to be compiled anywhere) on your computer.  But ... that uses your compiler, and compilers don't have to give exactly the same results.  And it's on your cpu, which doesn't have to be exactly the same as mine.  And it uses your computer's math libraries, which also don't have to be exactly the same as mine.

So all is lost?  Not really.  The thing is, these sorts of differences are small, and you can analyze your results with some care (and effort) to decide that the differences are indeed because compilers, processors, libraries, or operating systems, etc., don't always give exactly the same answers.  I've done this exercise before myself, as I was getting different answers from another group.  I eventually tracked down a true difference -- something beyond just the processors (etc.) doing things slightly differently (but different in a legal way).  The other group was doing some rounding in a way that I thought was incorrect, and which gave answers that differed from mine in the least significant bit about 3/4ths of the time.

With that understood, we could get exactly the same answers in all bits, all of the time, in spite of the different processors and such.  But, it was a lot of work to get to that point.  And this is a relatively simple system (in terms of the data and programs).

So are you lost on more complex systems, like general circulation models?  Again, no.  The thing is, if your goal is science -- understanding some part of nature -- you understand as well that computers aren't all identical and have done some looking in to how those differences affect your results.  The catchphrase here is "It's just butterflies".  Namely, the line goes that because weather is chaotic, a butterfly flapping its wings in Brazil today can lead to a tornado in Kansas five days from now.  What the catchphrase is referring to is that small differences (can't even call them errors, just different legal ways of interpreting or carrying out commands) in the computer can lead to observable changes down the road.  They don't change the main results -- if you're looking at climate, the onset of a tornado at 12:45 PM, April 3 1974 is not meaningfully different from 3:47 on the same day (though it certainly is if you live in the area!) -- but they do change some of the details.

What do we do, then?  At the moment, and I invite the software and hardware engineers out there to provide some education and corrections, what you need for exact reproducibility is to archive all the data, all the decoders, all the program source, all the compilers, all the system libraries, and all the hardware, exactly.  The full hardware archive is either impossible or close enough as makes no difference.  The system software archives (compilers and system libraries) are at least extraordinarily difficult and lie outside the hands of scientists (there's a reason they're called system tools).  The scientists' data and programs, not as easy as you might think, but doable.  Probably.

As you (I) then turn around and try to work with some archived processing system, then, when (not if) you get different results than the reference result, your first candidate for the difference is 'different operating system/system libraries/compilers/...', not dishonesty.  That means we have work to do, unfortunately.  I hope the software engineers have some good ideas/references/tools for making it easier.  I can say, though, from firsthand experience working with things I wrote 5-15 years ago, that there is just an astonishingly large number of ways you can get different answers -- even when you wrote everything yourself.  If you go on a fault-finding expedition, you'll find fault.  If you try to understand the science, though, even those 5-15 years worth of changes don't hide the science.

There's nothing special to data about this; models have the same issues.  Nor is there anything special to weather and climate.  The data and models used to construct a bridge, car, power plant, etc., also have these same issues.  If you've got some good answers to how to manage the issues, do contribute.  I'll add, though, that Steve (in his reply to my comments, a separate post, and a professional article) has found that climate models, from a software engineering perspective, appear to be higher quality than most commercial software.

Where is the surface?

I just commented on my facebook status that I'm at a meeting about sea surface temperature.  That part was safe.  Rest of the comment was to observe that I'm now back to wondering whether the sea has a surface, where it is if it does, and if it does, whether it has a temperature.  That prompted a friend to comment 'Great ... this is going to bug me now.'  So for him, here's a longer version.

This sort of question is very common to science.  Of course my musing for facebook is overstated.  But there is usually a real question about what exactly it is you've observed when you take an observation.  When you have very different observing methods, they may well observe things that are different from each other.  There are, let's say 4, different ways of observing the sea surface's temperature.  For a diagram, see the wikipedia article on sea surface temperature

The standard method, and reference for others, is calibrated buoys that carry a thermometer at a known depth, typically 1 meter.  A major drawback to this method (all methods of observing have drawbacks!) is that you need a buoy.  They're not cheap, and it would take several million of them to give us a high resolution data set for global sea surface temperature (acronymed SST).



The longest-used satellite method for obtaining SST is to observe the earth in infrared wavelengths.  Infrared doesn't propagate well through water, so the satellite sees a 'skin' temperature averages over 1 wavelength -- about 10 millionths of a meter (10 microns).  A drawback to this method is that it can't see through clouds.  Clouds also emit infrared, so the satellite tells you the temperature of the clouds, rather than the sea surface.  Weather forecasters use this to improve their forecasts.  But it means there's a lot of the earth these satellites can't see on any given day.

A recent addition to our satellite observing of sea surface temperature is to look at microwave wavelengths.  This usually can see through clouds as the wavelength is carefully chosen to be one that cloud drops don't emit or absorb much.  This gives us a temperature at about 1 mm depth.  The drawback for these is that the satellite averages over a large area -- 25-50 km diameter, as opposed to the 1-4 km for the infrared satellites.

The fourth class is the 'everything else' group: ship 'bucket' temperatures, hull contact temperatures, water intake temperatures, temperature sensors on chains (often from buoys) extending well below the surface (or 1 m depth), or the ARGO floats -- which observe temperatures from 2000 m depth up to close to the surface.  All of these observe temperatures at depths greater, sometimes much greater, than 1 meter below the surface.

If the 10 micron, 1 millimeter, 1 meter, 5-20 meter temperatures reported by the different methods were the same (within observing error), then it'd be a concern for specialists alone.  The reality, however, is that the different temperatures can be substantially different.  Most of the time, over most of the globe, they're close.  But once you have calm winds (less than 5 m/s, 10 mph, roughly) and strong sunlight, you can accumulate skin (that 10 micron temperature the infrared satellites see) heating enough to warm temperatures.  If the winds are very calm, under 2 m/s, it can be by a few degrees -- but only in the skin.  The 1 millimeter ('subskin') warms, but not by as much.  1 meter down the temperature may change only a little.  And at 5-20 meters, almost entirely unchanged.  So ... if you want the 'sea surface temperature', which of the 4 do you want?  And is that 5, or 20 meters for the 'unchanged'?  How close to entirely unchanged is close enough?

Hence my question: Where is the sea surface?

As mentioned by someone else -- under hurricane conditions, are you even sure that there is a sea surface?

How much detail is there really?

I'm thinking about sea surface temperature (SST) these days, but the approach here is one that can be applied to many situations, even ones outside weather and climate. A common, important, and not always easy, questions is -- just how much detail do you need? The more detail, the more expensive it is to make a good product, whether that's an analysis of sea surface temperature, a climate model, or a surface in a video game. Of course, what I'd like is the sea surface temperature every few meters over the entire globe. If that's more than necessary at some time, I could average it down. But ... it would take an awful lot of storage to save temperatures every few meters (my back yard, my neighbor's, my front yard, ...) over the whole globe.

Let's start by looking at an actual high resolution global product, though not every few meters! The SST analysis at http://polar.ncep.noaa.gov/sst/ gives a value every 1/12th of a degree in latitude and longitude, one about every 9 km (6 miles). It has about 9 million values. Let's also suppose that this is fine enough resolution that everything important is represented.

The worst resolution is to use 1 number for the entire globe, the average for all ocean points. To measure how bad this is, I'm going to compute the root mean square error. (Those who know what this is can skip to the next paragraph.) It is often abbreviated rmse. To find it, we go through every ocean point in the grid and find the difference between the value there and the average. Then we multiply this difference by itself (square it -- this avoids the marksmen statistician story*). Then add up these squares for every ocean point. This is a big and not interesting number. One thing that would be more interesting is the average value of the squared error -- the mean square error. So we divide by the number of points that were involved. This also tells us the error variance. Since we think more in terms of temperature and temperature changes than squares of temperature changes, we take the square root of the mean square error -- get the rmse. This is a figure which represents a typical magnitude of how far off we expect to be. We could be either warmer or colder by this much, but this is the magnitude.

* Two statisticians went to a shooting range and each fired at the target. The first missed by 1 meter to the left (-1 meter). The second missed by 1 meter to the right (+1 meter). They then congratulated each other on their fine marksmanship because on average they had hit the bullseye. Their average error was indeed zero. But their rms error was 1 meter.

When I compute the RMSE for using global mean temperature instead of the full resolution grid, I find 12 C. That's ... enormous. The difference between water at 20 C (68 F) and 32 C (90 F) is pretty large! So, clearly, we can't be satisfied with an RMSE of 12 C. But now we have a method for looking at the resolution we need, and a notion of how bad you can get.

Then I made my program average over smaller boxes than the whole globe, say 90 degrees on a side -- London to Chicago, equator to pole -- and found the RMSE comparing those box averages to the original temperatures in the full resolution grid. No surprise that boxes that large were pretty bad. But ... once I got down to boxes 2 degrees on a side (which is something like 200 km, or 120 miles), the RMSE was down to 0.5 degrees.

This is still definitely not zero, but it isn't bad. When a typical satellite used for the job -- such as the AVHRR instrument on NOAA-18 -- is used to make an observation, it has an RMSE (compared to a buoy's thermometer at about the same location at about the same time) of about 0.5 degrees. In other words, with boxes 2 degrees on a side, the average represents what is happening in sea surface temperature about as well as getting a single observation from satellite. We've also managed to reduce our RMS error by about 95% as compared to using only a single number. On the other hand, even though we've captured 95% of what's going on, we only need to use 16,200 numbers -- instead of the 9,331,200 we started with. 95% of the information of the full grid, with only 0.2% as much data.

We've caught 90% of what is happening (reduced the rmse by 90%) when the boxes are 6 degrees on a side (600 km, 360 miles). And it's 99% once we're down to boxes only 0.5 degrees (50 km, 30 miles) on a side (which means only about 3% as many data points are needed to represent the full data set to 99% accuracy).

Now, let's translate this back to some situations we might care about. In trying to construct climatologies of sea surface temperature, we run in to the problem that as we go back in time, there are fewer and fewer data points. On the other hand, if we have 1 observation in each box 6 degrees on a side, we've managed to capture 90% of what is happening in the sea surface temperature. In other words, a much sparser data set than we might imagine could indeed represent an awful lot of what is happening in the ocean. A global grid at 6 degrees resolution has only 1800 points, so we need only 1800 observations to fill it in our simple-minded way.

At 2 degrees resolution, we've captured 95% of what happens in sea surface temperature (at least to this quick little glance -- I only looked at 1 day, as analyzed by 1 center, etc.) So, if we had a good global ocean model at 2 degree resolution, we'd actually be pretty far along in being able to predict sea surface temperatures (model climate, etc.) well. In practice, there are processes that happen in smaller areas than the 2 degree box which can change the whole box's average and we, therefore, want finer resolution than 2 degrees. More about that in a different post.

In thinking about observing systems, if we only 'need' 1 observation every 200 km or so, and we have satellites that can take an observation every 4 km (like the one above) we're all done, right? Unfortunately, no. The problem is, that satellite looks for clouds. If there are clouds -- and cloudy areas can easily stretch for 1000 km -- the satellite can't see the sea surface to tell us what the temperature is down there. So we need other data sources -- ships, buoys, other sorts of satellite (ones that can see through clouds) to fill in even just the 200 km (2 degrees latitude-longitude) boxes each day. Plus we need to observe the detail in the oceans that are involved in those other processes I mentioned. It isn't just for models that they're important -- fishing also cares.
Older Post ►
eXTReMe Tracker
 

Copyright 2011 Grumbine Science is proudly powered by blogger.com