Notes from Ken Johnson
2005.02.15

What influences people in the way of tools?  ODV, LAS (more frustrated than happy), big things influencing things. MBSystem.  More at visualization end of things.  Not on data internal sides.

<KG>
Some other useful tools:
- Excel
- GMT
- His website: 134.89.11.194
</KG>

Big imperative?  Whole video thing, either stop collecting or figure out what to do with it.  Who's the big user to take advantage?

What worked well? 5 ways to slice: science, software, opportunities, community needs, what's been successful modes.

Big initiatives: Chemistry, biology, and low pH oceans; global nitrogen cycle; submarine mass movements; gas hydrates as carbon sink/source.  Of big MBARI initiatives, what are big science questions, then tunnel down to specific software you'd need.  Landslides -> MBSystem -> repeat mapping.  Hard to know what kinds of sensors if you don't know  (Ditto software?)   

<KG>
I got a sense in here (especially with the mass movements) is that getting science to define their needs is hard, because they don't always know what they need yet.  This was the idea behind the "scoping project", but I think we need to somehow account for the fact that it will always be REALLY hard to nail down requirements from science as they are not even sure what or how they need to get the information to answer the science questions.
</KG>

Role of individual PI vs institutional initiatives.  Group effort vs Eureka moment.  

Wishes for well outfitted ARGO float (GODAE hosts).  Want to link data into models.  Meld all these sources into one product.  What would this look like?  1 number of productivity.  

Quality control issues: drift, validation, volume.  

flux.ocean.washington.edu just gives single profiles. Cute but not useful.  you want to make maps of the anomaly. 894, 5 are examples.  use floats to build productivity model.  what you'd do for the PowerPoint is different than what you do to get the one number.

This is visualization, but what about queries?  Would like to be able to query for data.  Some of this is available.

Much more data with floats than ships.  Much in common with our local data.  How do we put them all into one format?  Would everyone want the same distillation?  3 packages are out there, underpinnings quite different (ODV, Ocean Atlas Viewer, Java Ocean Atlas (Same?).  

First thing before doing data is a product to see what is there.  Convincing customer that development is worth continuing.  Patch together to do visualization first.  Some things people really want will become evident only when visualization is available.

DE: Biggest weakness is size.  Where is biggest impact?  Another way to look at that is how big are other groups that are operating?  ODV: 1 person.  MBSystem: 1 person.  These are long-term projects with one focused person. MBSystem, ODV developers wrote them to satisfy themselves.

Moving on to other scenarios...low pH ocean.  Don't have good instruments, or mechanisms for measuring.  Envision changes on organisms.  FRRF to see how plants respond, benthic respirometer or eddy flux correlation instruments.  Very noisy measurement.

Sediment transport: sensorizing the bottom.  Modeling components? Many types.  We know so little.  Societal significance.

Gas hydrates.  (Science overview.)

Effort at merging platforms is common data management is a common aspect to all these.  [Note from John: How would this be used?  What about designing an interface between models and observing systems that lets models ask for observations meeting particular parameters.] 

Take ODV and go one step further, not just variable-variable but correlations, cycles.  Imagine 40-mooring array and think about measuring transport -- with enough floats you can measure the entire heat content, you don't care about transport.  With 40 you've resolved the spatial.  John Ryan would ask different questions.

JB: Different scientists do plots from same data differently.  Is there really a tool for all?  ODV doesn't constrain user too far, many capabilities for scripting etc.  How is it successful with such a hard interface?  Show capabilities with a few built-in examples that are easy.  Demonstrate value early.

Important to have bio-informatic side of things.  Haddock/Vrijenhoek/Scholin to talk.  Ed needed more connection to bio-informatics people.  Mapping data were here (or we're here).  Will always be an area for something like ReefGrow.  Is this an engineering task, or a scientists' staff?  Can't be cost free to scientists.  Don't do a good job of allocating costs back to program.  Tragedy of the commons (burning up common resources) can result.

Are existing systems relevant?  SSDS, AVED.  Looks at ODV files to access SSDS data.  Direct use is ISUS data.  When we make ODV file.  Mostly deal with M1/M2 data.  

What about the opportunities list?  Pick any tool and you could use it as a framework for addressing these.  These bullets are pretty generic.  They are capabilities. Jim B: Look at it as "why wouldn't we use archival data?"  Many of the bullets hit that.  Processes are not scalable.  

Two issues with data quality control: 1) Automating it (setting up rules, pretty simple set).  2) Human eyeballs ("this other rule we forgot to tell you about").  ISUS now QC'd automatically.  Works because it's a complicated measurement, if you can't fit the array to a model you just toss it out.  Temperature is harder.  

Jim B: Lots of problems pop out when combining data sets.  Survey result can be internally consistent, globally misplaced -- can't tell until combining.  This is argument for combining lots of data sets. Lots of little things can be done to improve QC.  Sometimes a little thing can be very important to oceanography.  In astronomy data quality is designed into instruments (but there are reasons for that, like sensitivity and extremity reasons).

What are other science questions?  We got more process issues that affect multiple science questions.

What NSF ought to do is set up their own data servers, at data.nsf.org, where people could submit their data and point to it later on, let it be google'd, etc.

