HR: 1340h
AN: NG43A-0439    [Abstracts]
TI: How to Leverage Existing Spectral Knowledge when Clustering Hyperspectral Data
AU: * Wagstaff, K
EM: kiri.wagstaff@jpl.nasa.gov
AF: Jet Propulsion Laboratory, 4800 Oak Grove Drive, Pasadena, CA 91109 United States
AU: Shu, H
EM: sung@its.caltech.edu
AF: California Institute of Technology, 1200 East California Blvd., Pasadena, CA 91125 United States
AU: Castano, R
EM: rebecca.castano@jpl.nasa.gov
AF: Jet Propulsion Laboratory, 4800 Oak Grove Drive, Pasadena, CA 91109 United States
AB: Hyperspectral images collect large volumes of data, with observations at hundreds or thousands of different wavelengths. The large data size renders a thorough manual analysis difficult, expensive, and time-consuming. For example, the Hyperion instrument on the EO-1 spacecraft regularly produces image cubes that are over 1 gigabyte in size (256 $\times$ 7000 pixels, at 242 wavelengths). Automated techniques for analyzing and summarizing these mega-data sets provide two major benefits: 1) scientists can quickly obtain high-level views of the data contents, and 2) summaries produced on-board the spacecraft enable quick prioritization of data for transmission to make the best use of limited bandwidth. One approach for generating summaries is to partition the pixels from a given image into a set of $k$ clusters. Each cluster contains pixels that are more similar to each other than to pixels in other clusters. The image can be summarized by the set of $k$ clusters, represented by the cluster means and standard deviations (of the pixel values for each cluster). However, typical clustering algorithms are completely data-driven and will produce summaries based on the largest distinguishing factor between pixels (often, brightness), regardless of whether that distinction is physically meaningful. In contrast, we seek to include knowledge from existing spectral libraries to improve the automated summaries. In this work, we present a knowledge-driven clustering method that incorporates laboratory spectra as ``seeds'' for some, or all, of the data clusters. We contrast the summaries produced by data-driven and knowledge-driven clustering. We find that summaries that incorporate the spectral library have greater science value than those produced from the data alone. Knowledge-based summaries are more interpretable and more likely to be based on true compositional differences in the areas being imaged. We present sample results from several diverse areas to illustrate the benefits of clustering with spectral libraries.
DE: 5464 Remote sensing
DE: 5494 Instruments and techniques
SC: Nonlinear Geophysics [NG]
MN: 2004 AGU Fall Meeting