HR: 14:25h
AN: IN33D-04 [Abstracts]
TI: A Testbed Demonstration of an Intelligent Archive in a Knowledge Building System
AU: * Ramapriyan, H
EM: Rama.Ramapriyan@nasa.gov
AF: NASA, Goddard Space Flight Center, Greenbelt, MD 20771
United States
AU: Isaac, D
EM: david.isaac@teambps.com
AF: Business Performance Systems, 6911 Fairfax Road, Bethesda, MD 20814
United States
AU: Morse, S
EM: smorse@sosacorp.com
AF: SoSACorp, 3877 Fairfax Ridge Road, Ste. 201-C, Fairfax, VA 22030-7425
United States
AU: Yang, W
EM: yang@rattler-e.gsfc.nasa.gov
AF: Laboratory for Advanced Information Technology and Standards, George Mason University, 6301 Ivy Lane,
Suite 620, Greenbelt, MD 20770
United States
AU: Bonnlander, B
EM: bbonnlander@ihmc.us
AF: Institute for Human and Machine Cognition, 40 S. Alcaniz St, Pensacola, FL 32501
United States
AU: McConaughy, G
EM: Gail.R.McConaughy@nasa.gov
AF: NASA, Goddard Space Flight Center, Greenbelt, MD 20771
United States
AU: Di, L
EM: ldi@gmu.edu
AF: Laboratory for Advanced Information Technology and Standards, George Mason University, 6301 Ivy Lane,
Suite 620, Greenbelt, MD 20770
United States
AU: Danks, D
EM: ddanks@andrew.cmu.edu
AF: Department of Philosophy, Carnegie Mellon University, 135 Baker Hall, Pittsburgh, PA 15213
United States
AB:
The last decade's influx of raw data and derived geophysical parameters from several Earth observing satellites to NASA data
centers has created a data-rich environment for Earth science research and applications.ÿ While advances in hardware and
information management have made it possible to archive petabytes of data and distribute terabytes of data daily to a broad
community of users, further progress is necessary in the transformation of data into information, and information into
knowledge that can be used in particular applications in order to realize the full potential of these valuable datasets.
In examining what is needed to enable this progress in the data provider environment that exists today and is expected to
evolve in the next several years, we arrived at the concept of an Intelligent Archive in context of a Knowledge Building
System (IA/KBS). Our prior work and associated papers investigated usage scenarios, required capabilities, system
architecture, data volume issues, and supporting technologies. We identified six key capabilities of an IA/KBS: Virtual
Product Generation, Significant Event Detection, Automated Data Quality Assessment, Large-Scale Data Mining, Dynamic Feedback
Loop, and Data Discovery and Efficient Requesting.
Among these capabilities, large-scale data mining is perceived by many in the community to be an area of technical risk. One
of the main reasons for this is that standard data mining research and algorithms operate on datasets that are several
orders of magnitude smaller than the actual sizes of datasets maintained by realistic earth science data archives. Therefore,
we defined a test-bed activity to implement a large-scale data mining algorithm in a pseudo-operational scale environment
and to examine any issues involved. The application chosen for applying the data mining algorithm is wildfire prediction
over the continental U.S. This paper reports a number of observations based on our experience with this test-bed.
While proof-of-concept for data mining scalability and utility has been a major goal for the research reported here, it was
not the only one. The other five capabilities of an IA/KBS named above have been considered as well, and an assessment of
the implications of our experience for these other areas will also be presented. The lessons learned through the testbed
effort and presented in this paper will benefit technologists, scientists, and system operators as they consider introducing
IA/KBS capabilities into production systems.
DE: 0520 Data analysis: algorithms and implementation
DE: 0525 Data management
DE: 0530 Data presentation and visualization
DE: 0535 Hardware solutions
DE: 0555 Neural networks, fuzzy logic, machine learning
SC: Earth and Space Science Informatics [IN]
MN: Fall Meeting 2005