IN53C-01
Evolution of Metadata Standards: New Features in ISO 19115
Metadata standards in the earth sciences have evolved considerably over the last several decades. Drivers of this evolution have included community needs initially expressed as extensions and then incorporated into future standards, as well as the recognition of the importance of preserving understanding relative to data discovery. The ISO Geographic Metadata Standard, 19115, includes many features that originally emerged in the Directory Interchange Format (DIF) and as extensions to the FGDC Metadata Standards. The ISO Standard also includes new and powerful features that extend the capabilities of the standard considerably beyond previous standards. These features include a spectrum of metadata from granule to collection levels, user feedback, reasonable on- line resources, multiple spatial and temporal extents, and a very general mechanism for describing data quality. I will describe these features with examples of how they will be used at NGDC.
IN53C-02
GeochronML - O&M, GeoSciML and the Age of the Earth: Developing a geochronology data model
Geochronology is the vital fourth dimension for geological knowledge. It provides the temporal framework for understanding and modelling geological processes and rates of change. Incorporating geochronological ‘observations and measurements' into interoperable geological data systems is thus a critical pursuit. Although there are several resources for storing and accessing geochronological data, there is no standard format for exchanging such data among users. Current systems are a mixture of comma-delimited text files, Excel spreadsheets and PDFs that assume prior specialist knowledge and frequently force the user to laboriously, and potentially erroneously, extract the required data manually. Geoscience Australia and partners are developing a standard data exchange format for geochronological data (`geochronML') within the broader framework of Observations and Measurements and GeoSciML that are an important facet of emerging international geoscience data transfer standards. Geochronology analytical processes and resulting data present some challenging issues as a rock `age' is typically not a direct measurement, but rather the interpretation of a statistical amalgam of several measurements chosen with the aid of prior geological knowledge and analytical metadata. The level at which these data need to be exposed to a user varies greatly, even to the same user over the course of a project. GeochronML is also attempting to provide a generic pattern that will support as wide as range of radioisotopic systems as possible. This presentation will discuss developments at Geoscience Australia and the opportunities for collaboration. http://www.seegrid.csiro.au/twiki/bin/view/AnalyticalGeoscience/Geochronology
IN53C-03
GroundWater Markup Language (GWML): Extending GeoSciML for Groundwater
The emerging growth of OGC web services in the groundwater domain is likely to follow trends in other disciplines and result in a rich supply of groundwater data in heterogeneous and complex GML formats (Geography Markup Language), even within a single organization. The Geological Survey of Canada (GSC) has developed the Groundwater Markup Language (GWML) to overcome barriers to the use of the data in such environments. GWML is a common format for exchanging groundwater data. It extends two advanced GML standards, GeoSciML and Observations & Measurements, by adding entities such as hydrogeologic units (e.g. aquifers), properties (e.g. storativity), water wells, and water budget entities. Due to this firm grounding in OGC standards, GWML can be used with several OGC services such as WFS, WMS, WCS, and SOS, to enable exchange of a wide spectrum of groundwater data. The design, development and intended use of GWML will be discussed, including its relation to the other GML standards and its role as an integral part of GSC's emerging groundwater information system.
IN53C-04
When Metadata Alone Is Not Enough
The Virtual ITM Observatory (VITMO) is currently operational and provides unique search and data delivery capabilities to the Ionosphere Thermosphere Mesosphere (ITM) community. Many of these capabilities in data search are currently constrained by the lack of detailed metadata required to support user requests. For example, determining when a remote sensing instrument observes a ground site is not knowable using today's metadata. In VITMO, we enhance our search capabilities by tying together the standard SPASE compliant metadata, virtual metadata generators, solar geophysical indices databases, and event lists. These additional databases and tools provide the enhanced capabilities that allow the VITMO to answer detailed searches across a variety of data sources. We will detail the architecture of this system and show how these capabilities can allow the researcher to perform many levels of scientific analysis directly from the search system. VITMO can be found at http://vitmo.jhuapl.edu. http://vitmo.jhuapl.edu
IN53C-05
Towards Long-Term Archiving of NASA HDF-EOS and HDF Data - Data Maps and the Use of Mark-Up language
The Hierarchical Data Format (HDF) has been a data format standard in NASA's Earth Observing System Data and Information System (EOSDIS) since the 1990s. Its rich structure, platform independence, full-featured Application Programming Interface (API), and internal compression make it very useful for archiving science data and utilizing them with a rich set of software tools. However, a key drawback for long-term archiving is the complex internal byte layout of HDF files, requiring one to use the API to access HDF data. This makes the long-term readability of HDF data for a given version dependent on long-term allocation of resources to support that version. The majority of the data from NASA's Earth Observing System (EOS) have been archived in HDF Version 4 (HDF4) format. To address the long-term archival issues for these data a collaborative study between The HDF Group and NASAs EOSDIS data centers is underway. One of the first activities undertaken has been an assessment of the range of HDF4 formatted data held by NASA to determine the capabilities inherent in the HDF format that have been used in practice. Based on the results of this assessment, methods for producing a map of the layout of the HDF Version 4 files held by NASA will be prototyped using a markup-language-based HDF tool to map the layout of the HDF Version 4 files. The resulting maps should allow a separate program to read the file without recourse to the HDF API. To verify this, two independent tools based solely on the map files will be developed and tested with a variety of data products archived by NASA.
IN53C-06
Metadata Management on the SCEC PetaSHA Project: Helping Users Describe, Discover, Understand, and Use Simulation Data in a Large-Scale Scientific Collaboration
Large scientific collaborations, such as the SCEC Petascale Cyberfacility for Physics-based Seismic Hazard Analysis (PetaSHA) Project, involve interactions between many scientists who exchange ideas and research results. These groups must organize, manage, and make accessible their community materials of observational data, derivative (research) results, computational products, and community software. The integration of scientific workflows as a paradigm to solve complex computations provides advantages of efficiency, reliability, repeatability, choices, and ease of use. The underlying resource needed for a scientific workflow to function and create discoverable and exchangeable products is the construction, tracking, and preservation of metadata. In the scientific workflow environment there is a two-tier structure of metadata. Workflow-level metadata and provenance describe operational steps, identity of resources, execution status, and product locations and names. Domain-level metadata essentially define the scientific meaning of data, codes and products. To a large degree the metadata at these two levels are separate. However, between these two levels is a subset of metadata produced at one level but is needed by the other. This crossover metadata suggests that some commonality in metadata handling is needed. SCEC researchers are collaborating with computer scientists at SDSC, the USC Information Sciences Institute, and Carnegie Mellon Univ. in order to perform earthquake science using high-performance computational resources. A primary objective of the "PetaSHA" collaboration is to perform physics-based estimations of strong ground motion associated with real and hypothetical earthquakes located within Southern California. Construction of 3D earth models, earthquake representations, and numerical simulation of seismic waves are key components of these estimations. Scientific workflows are used to orchestrate the sequences of scientific tasks and to access distributed computational facilities such as the NSF TeraGrid. Different types of metadata are produced and captured within the scientific workflows. One workflow within PetaSHA ("Earthworks") performs a linear sequence of tasks with workflow and seismological metadata preserved. Downstream scientific codes ingest these metadata produced by upstream codes. The seismological metadata uses attribute-value pairing in plain text; an identified need is to use more advanced handling methods. Another workflow system within PetaSHA ("Cybershake") involves several complex workflows in order to perform statistical analysis of ground shaking due to thousands of hypothetical but plausible earthquakes. Metadata management has been challenging due to its construction around a number of legacy scientific codes. We describe difficulties arising in the scientific workflow due to the lack of this metadata and suggest corrective steps, which in some cases include the cultural shift of domain science programmers coding for metadata.
IN53C-07
Data Publishing and Sharing Via the THREDDS Data Repository
The terms "Team Science" and "Networked Science" have been coined to describe a virtual organization of researchers tied via some intellectual challenge, but often located in different organizations and locations. A critical component to these endeavors is publishing and sharing of content, including scientific data. Imagine pointing your web browser to a web page that interactively lets you upload data and metadata to a repository residing on a remote server, which can then be accessed by others in a secure fasion via the web. While any content can be added to this repository, it is designed particularly for storing and sharing scientific data and metadata. Server support includes uploading of data files that can subsequently be subsetted, aggregrated, and served in NetCDF or other scientific data formats. Metadata can be associated with the data and interactively edited. The THREDDS Data Repository (TDR) is a server that provides client initiated, on demand, location transparent storage for data of any type that can then be served by the THREDDS Data Server (TDS). The TDR provides functionality to: \item* securely store and "own" data files and associated metadata \item* upload files via HTTP and gridftp \item* upload a collection of data as single file \item* modify and restructure repository contents \item* incorporate metadata provided by the user \item* generate additional metadata programmatically \item* edit individual metadata elements The TDR can exist separately from a TDS, serving content via HTTP. Also, it can work in conjunction with the TDS, which includes functionality to provide: \item* access to data in a variety of formats via
IN53C-08
Design and implementation of CUAHSI WaterML and WaterOneFlow Web Services
WaterOneFlow is a term for a group of web services created by and for the Consortium of Universities for the Advancement of Hydrologic Science, Inc. (CUAHSI) community. CUAHSI web services facilitate the retrieval of hydrologic observations information from online data sources using the SOAP protocol. CUAHSI Water Markup Language (below referred to as WaterML) is an XML schema defining the format of messages returned by the WaterOneFlow web services. \newline WaterML CONCEPTS \newline The goal of the first version of WaterML was to encode the semantics of hydrologic observations discovery and retrieval and implement WaterOneFlow services in a way that creates the least barriers for adoption by the hydrologic research community. In particular, this implied maintaining a single common representation for the key constructs returned on web service calls, related to observations, features of interest, observation procedures, observation series, etc. In addition to point measurements described in the CUAHSI Observations Database Model (ODM) specification, hydrologic information may be available as observations or model outcomes aggregated over user-defined regions or grid cells. While USGS NWIS and EPA STORET exemplify the former case, sources such as MODIS and Daymet are examples of the latter. In this latter case, as in the case of other remote sensing products or model-generated grids, the observation or model-generated data are treated as fields, and sources of such data are referenced in WaterML as datasets, as opposed to sites. \newline IMPLEMENTATION \newline WaterML is primarily designed for relaying fundamental hydrologic time series data and metadata between clients and servers, and to be generic across different data providers. Different implementations of WaterOneFlow services may add supplemental information to the content of messages. However, regardless of whether or not a given WaterML document includes supplemental information, the client shall be sure that the portion of WaterML pertaining to space, time, and variables will be consistent across any data source. Depending on the type of information that the client requested, a WaterOneFlow web service will assemble the appropriate XML elements into a WaterML response, and deliver that to the client. The core WaterOneFlow methods include: \newline * GetSiteInfo – for requesting information about an observations site, and returns site properties, and observation series for that site. \newline * GetVariableInfo – for requesting information about a variable, and returns variable properties. \newline * GetValues – for requesting a time series for a variable at a given site or spatial fragment of a dataset, and returns data values, and associated variable and site properties to make a single self-contained request. \newline The CUAHSI WaterML description has been submitted as a discussion paper to the Open Geospatial Consortium, and is available from the OGC portal. The project web site is http://www.cuahsi.org/his . http://www.cuahsi.org/his