HR: 14:25h
AN: IN33D-04    [Abstracts]
TI: A Testbed Demonstration of an Intelligent Archive in a Knowledge Building System
AU: * Ramapriyan, H
EM: Rama.Ramapriyan@nasa.gov
AF: NASA, Goddard Space Flight Center, Greenbelt, MD 20771 United States
AU: Isaac, D
EM: david.isaac@teambps.com
AF: Business Performance Systems, 6911 Fairfax Road, Bethesda, MD 20814 United States
AU: Morse, S
EM: smorse@sosacorp.com
AF: SoSACorp, 3877 Fairfax Ridge Road, Ste. 201-C, Fairfax, VA 22030-7425 United States
AU: Yang, W
EM: yang@rattler-e.gsfc.nasa.gov
AF: Laboratory for Advanced Information Technology and Standards, George Mason University, 6301 Ivy Lane, Suite 620, Greenbelt, MD 20770 United States
AU: Bonnlander, B
EM: bbonnlander@ihmc.us
AF: Institute for Human and Machine Cognition, 40 S. Alcaniz St, Pensacola, FL 32501 United States
AU: McConaughy, G
EM: Gail.R.McConaughy@nasa.gov
AF: NASA, Goddard Space Flight Center, Greenbelt, MD 20771 United States
AU: Di, L
EM: ldi@gmu.edu
AF: Laboratory for Advanced Information Technology and Standards, George Mason University, 6301 Ivy Lane, Suite 620, Greenbelt, MD 20770 United States
AU: Danks, D
EM: ddanks@andrew.cmu.edu
AF: Department of Philosophy, Carnegie Mellon University, 135 Baker Hall, Pittsburgh, PA 15213 United States
AB: The last decade's influx of raw data and derived geophysical parameters from several Earth observing satellites to NASA data centers has created a data-rich environment for Earth science research and applications.ÿ While advances in hardware and information management have made it possible to archive petabytes of data and distribute terabytes of data daily to a broad community of users, further progress is necessary in the transformation of data into information, and information into knowledge that can be used in particular applications in order to realize the full potential of these valuable datasets. In examining what is needed to enable this progress in the data provider environment that exists today and is expected to evolve in the next several years, we arrived at the concept of an Intelligent Archive in context of a Knowledge Building System (IA/KBS). Our prior work and associated papers investigated usage scenarios, required capabilities, system architecture, data volume issues, and supporting technologies. We identified six key capabilities of an IA/KBS: Virtual Product Generation, Significant Event Detection, Automated Data Quality Assessment, Large-Scale Data Mining, Dynamic Feedback Loop, and Data Discovery and Efficient Requesting. Among these capabilities, large-scale data mining is perceived by many in the community to be an area of technical risk. One of the main reasons for this is that standard data mining research and algorithms operate on datasets that are several orders of magnitude smaller than the actual sizes of datasets maintained by realistic earth science data archives. Therefore, we defined a test-bed activity to implement a large-scale data mining algorithm in a pseudo-operational scale environment and to examine any issues involved. The application chosen for applying the data mining algorithm is wildfire prediction over the continental U.S. This paper reports a number of observations based on our experience with this test-bed. While proof-of-concept for data mining scalability and utility has been a major goal for the research reported here, it was not the only one. The other five capabilities of an IA/KBS named above have been considered as well, and an assessment of the implications of our experience for these other areas will also be presented. The lessons learned through the testbed effort and presented in this paper will benefit technologists, scientists, and system operators as they consider introducing IA/KBS capabilities into production systems.
DE: 0520 Data analysis: algorithms and implementation
DE: 0525 Data management
DE: 0530 Data presentation and visualization
DE: 0535 Hardware solutions
DE: 0555 Neural networks, fuzzy logic, machine learning
SC: Earth and Space Science Informatics [IN]
MN: Fall Meeting 2005