IN11B-0460
A Tool for Distributed Inquiry-Based Exploration of MHD Model Output
Adaptive grid models that can simulate the magnetohydrodynamic nature of solar storms and their propagation through space and interaction with the Earth's magnetic fields requires significant expertise and computer resources. Consequently the number of undergraduate and graduate students who will have access to opportunities to work with these models and their output is necessarily limited. This paper reports on an effort to create a tool that can be used to display MHD model results in a 3D visualization connected to a notation system so that students remote from centers of excellence in MHD modeling can work with results as part of class exercises. The goal is to expand the opportunities for MHD model research to a broader range of students and faculty than is possible today. http://www.mhdmodel.org
IN11B-0461
MLSgrid: Testing and Deployment of a Grid Network for Science Data Analysis
The Microwave Limb Sounder (MLS) Science Computing Facility (SCF) stores over 50 terabytes of data, has over 240 computer processing hosts, and 64 users from around the world. These resources are spread over three primary geographical locations - the Jet Propulsion Laboratory (JPL), Raytheon RIS, and New Mexico Institute of Mining and Technology (NMT). A need for a grid network system was identified and defined to solve the problem of users competing for finite, and increasingly scarce, MLS SCF computing resources. Using Sun's Grid Engine software, a grid network was successfully created in a development environment that connected the JPL and Raytheon sites, established master and slave hosts, and demonstrated that transfer queues for jobs can work among multiple clusters in the same grid network. This poster will first describe MLS SCF resources and the lessons that were learned in the design and development phase of this project. It will then go on to discuss the test environment and plans for deployment by highlighting benchmarks and user experiences.
IN11B-0462
Using Grid Technologies to Encourage Web Service Developement in Heliophysics Virtual Observatories
The Heliophysics data environment is undergoing rapid expansion with the emergence of Virtual Observatories (VOs). These VOs allow transparent access to distributed data sets via a standard search interface. A key anticipated component of the data environment and VOs is the addition of web services to carry out additional data processing and visualization after a user has found data of interest with a VO. While VOs are reaching their prototype stages the emergence of services has been limited. We explore a novel approach to service developement and maintanance that uses Grid technologies. The system allows the community to easily turn processing code into a web service and grants them the use of spare processing power and storage space for service execution. Additionally, we emphasize the benefits afforded to the user community from the use of Grid technologies.
IN11B-0463
Pegasus: Providing Computation Management for Earth Science Applications
Earth science applications such as those being developed within the Southern California Earthquake Center (SCEC) require the coordination of hundreds of thousands of computation and data management tasks running on the national cyberinfrastructure resources such as the TeraGrid. Such coordination cannot be done by a single person; rather, automated techniques such as those provided by the Pegasus workflow management system need to be put in place. In this talk we will describe the Pegasus software from the point of view of an application and show examples of use of Pegasus in SCEC applications such as CyberShake and Earthworks. The SCEC CyberShake Project uses physics-based models of earthquake processes and integrates these models into a scientific framework for seismic hazard analysis and risk management. As a result CyberShake aims to produce more accurate hazard curves for the Southern California area. These hazard curves can then be interpolated to produce a Probabilistic Seismic Hazard Analysis (PSHA) map for a given area. Probabilistic seismic hazard maps indicate the likelihood of seeing a specific amount of surface motion within a specified period of time. PSHA is used (among other things) to determine the placement and design of buildings and other structures. The SCEC Earthworks Science Gateway is designed to compute and distribute ground-motion simulations for use in risk assessment and earthquake-engineering analysis. Earthworks uses high-performance simulation software called Anelastic Wave Propagation (AWP) models. AWP's compute the propagation, interference, and attenuation of seismic waves as they travel from a fault rupture to a target site. The results are typically vector-valued ground velocity values as a function of time, from which essentially any intensity measure can be computed. Both CyberShake and Earthworks require high-performance, multi-processor resources such as those provided by the NSF-funded TeraGrid to perform the computations in a reasonable amount of time, such as in hours instead of weeks. However, in order to use these resources easily and efficiently, tools such as Pegasus, Condor DAGMan, and grid-based services such as GridFTP need to be used. Pegasus is designed to help scientists describe a series of scientific computations (a workflow) in abstract terms without worrying about the details of the underlying infrastructure. These initial workflow descriptions are abstract, by which we mean that the workflows do not identify the particular resources needed for execution. The selection of specific resources to use is automatically performed by Pegasus which selects which resources to use, makes sure that the data is delivered to the computations, and delivers the results to the scientist. As a result Pegasus generates a concrete workflow, that is, an executable workflow which is executed reliably on the TeraGrid resources. Pegasus and DAGMan are not application or execution environment-specific. They have been applied in a variety of application from astronomy, gravitational-wave physics, neuroscience, and others running on national cyberinfrastructure as well as campus clusters, Condor pools, and desktops. http://pegasus.isi.edu/
IN11B-0464
3-D numerical seismic stratigraphic model generated from a highly anisotropic and wide- meshed survey grid.
Recently, 55 high-resolution seismic sections were collected by the Geological Survey of Canada to map the Quaternary sedimentary succession over an area of 7600 km2 in the St. Lawrence Estuary (eastern Canada). To better understand the geometrical relationships between these various units and to document the impact of the bedrock topography on the Quaternary basin infill, a numerical seismic stratigraphic model was developed. The main challenges related to its realization were the highly anisotropic character of the seismic grid due to the N330 orientation of most sections and the spacing between the sections ranging from 2.5 to 10 km. On each section, horizon picking of the key seismic unit boundaries generated a point coverage (x,y,z) that was later converted into curves. The dataset was sampled to increase the ratio between the distance of the neighbouring points and the distance between the curves. Control curves parallel to the basin axis were also created to constrain major topographic features to be modelled. Then, preliminary surfaces representing the superior limit of the seismic units were built with a discrete smooth interpolator by using the curves. This interpolator was robust enough to link the curves together with a minimum global roughness even if they were separated by a significant distance. The original point coverage was subsequently used as control nodes that forced the preliminary surfaces to be connected to the points. The influence of the wide-meshed character of the survey grid on the quality of the surface rendering was also corrected by using an equilateral triangulated mesh that included as many equilateral triangles as possible. Moreover this step made the surfaces more stable for numerical calculations. The resulting surfaces were used to create a volume model of 11 520 km3 that has horizontal and vertical resolutions of 250 and 5 m respectively. This model reflects adequately first order features such as dip and thickness variations of seismic units. It is also in agreement with more subtle features such as discontinuous bedrock highs that have been independently imaged on multibeam bathymetry. Depth slices and cross-sections can be easily extracted from the 3-D model and volume calculations can be done to improve geological interpretations in the area. The workflow used represents an effective way to generate a numerical model at the basin scale based on data collected according to a highly anisotropic and wide-meshed survey grid.
IN11B-0465
The Remote NetCDF Invocation (RNI) middleware platform. Making Scientific Datasets Available for Ubiquitous Computing.
Large holding of NetCDF data, such as in the Earth System Grid (ESG) or the Community Spectro-Polarimetric Analysis Center (CSAC) are vast repositories of data, making it if not impossible, but impractical for users to download and replicate the complete database. Furthermore, each individual dataset is a combination of hundreds of individual NetCDF files. Therefore requesting such dataset for analysis is an expensive transaction for individuals seeking ubiquitous computing. Since the current state of networks can provide for access to individual pieces of the dataset with enough reliability and speed, we seek a solution that will avoid the bulk download of the dataset required a priori, and will instead request needed portions of the dataset just-in-time. In order to achieve this, we modify the NetCDF C library to execute Remote NetCDF Invocation (RNI), that is, to operate on remote dataset, over HTTPS and gsiFTP protocols, individual NetCDF Application Programming Interface (API) calls as if they were local. This mechanism resembles the well known Remote Procedure Call (RPC) yet it radically differs on the binding between local and remote operations. Our design is based on the extensibility mechanism provided by the popular OPeNDAP Back-End Server (BES) middleware platform with Globus GridFTP and Apache modules acting as the proxy transport mechanism (binding) between the local and remote transactions. This paper describes the architecture as well as how we address the technical challenges for the complete system.
IN11B-0466
Parallel Analysis of Spectro-Polarimetric Signals in Heterogeneous Grids, Using OPeNDAP BES to Perform Scatter-Gather High Performance Computing.
Determination of the magnetic field of the Sun's photosphere from spectral images requires a complex mathematical process that translates into intense computational effort. Since each individual pixel of the image is independent of all others, this problem is well suited for scatter-gather parallel computing. At the High Altitude Observatory, we have constructed a dedicated Grid with the help of OPeNDAP version 4, also known as Back-End-Server (BES), as the middleware that facilitates all the interprocess communication required to find the complete cohesive solution for each dataset. This paper describes in detail the parallel approach taken to speed up the computation of the solution, as well as the specifics of the Grid design.
IN11B-0467
Automating Visualization Service Generation with the WATT Compiler
As tasks and workflows become increasingly complex, software developers are devoting increasing attention to automation tools. Among many examples, the Automator tool from Apple collects components of a workflow into a single script, with very little effort on the part of the user. Tasks are most often described as a series of instructions. The granularity of the tasks dictates the tools to use. Compilers translate fine-grained instructions to assembler code, while scripting languages (ruby, perl) are used to describe a series of tasks at a higher level. Compilers can also be viewed as transformational tools: a cross-compiler can translate executable code written on one computer to assembler code understood on another, while transformational tools can translate from one high-level language to another. We are interested in creating visualization web services automatically, starting from stand-alone VTK (Visualization Toolkit) code written in Tcl. To this end, using the OCaml programming language, we have developed a compiler that translates Tcl into C++, including all the stubs, classes and methods to interface with gSOAP, a C++ implementation of the Soap 1.1/1.2 protocols. This compiler, referred to as the Web Automation and Translation Toolkit (WATT), is the first step towards automated creation of specialized visualization web services without input from the user. The WATT compiler seeks to automate all aspects of web service generation, including the transport layer, the division of labor and the details related to interface generation. The WATT compiler is part of ongoing efforts within the NSF funded VLab consortium [1] to facilitate and automate time-consuming tasks for the science related to understanding planetary materials. Through examples of services produced by WATT for the VLab portal, we will illustrate features, limitations and the improvements necessary to achieve the ultimate goal of complete and transparent automation in the generation of web services. In particular, we will detail the generation of a charge density visualization service applicable to output from the quantum calculations of the VLab computation workflows, plus another service for mantle convection visualization. We also discuss WATT-LIVE [2], a web-based interface that allows users to interact with WATT. With WATT-LIVE users submit Tcl code, retrieve its C++ translation with various files and scripts necessary to locally install the tailor-made web service, or launch the service for a limited session on our test server. This work is supported by NSF through the ITR grant NSF-0426867. [1] Virtual Laboratory for Earth and Planetary Materials, http://vlab.msi.umn.edu, September 2007. [2] WATT-LIVE website, http://vlab2.scs.fsu.edu/watt-live, September 2007.
IN11B-0468
Toolkits for Automatic Service Generation: WATT and Kill-A-WATT
As part of the NSF funded VLab consortium [1], we have been involved in the automatic generation of visualization web services using the Web Automation and Translation Toolkit (WATT) compiler. The WATT compiler converts VTK Tcl input scripts into equivalent yet more efficient C++ web services by interpreting code structure, translating and then integrating bindings to the gSOAP library. WATT seeks to completely automate code distribution, integration of transport protocols and interface generation. Ideally, developers should concentrate on writing core applications, and let WATT transform them into web services in the background. Currently, the WATT compiler is limited to converting known Tcl commands and types to C++. For VTK a simple one to one mapping between Tcl and C++ is enforced, but Tcl commands without direct mappings slow the compilation process and require new mappings to be created. Loops and conditional statements are not yet implemented. In an effort to move forward with automation and not get caught up in the details of cross-language compilation, we developed a new application: Kill-A-WATT (KWATT). KWATT is a C++ application that utilizes the C++/Tcl library [2] to evaluate Tcl input scripts using the official Tcl interpreter. During evaluation of the input script, KWATT interprets code structure, integrating communication details via a Tcl-specific SOAP library [3]. Since KWATT drives the Tcl interpreter, the application has access to the full Tcl command base plus the ability to load new commands from other packages. KWATT is not a compiler; instead, it is a stand-alone application that is itself a web service. When KWATT consumes Tcl input, the generated web methods extend the list of previously available commands. This implies that C++ web methods statically defined in KWATT provide a set of standard methods available to every service. Also, since KWATT uses the Tcl interpreter, it has the potential to accept additional Tcl at any time while running, thereby allowing for patches or updates to the running service without downtime. We will discuss the current development status of KWATT (work in progress), and some of its future applications. In particular we will provide a comparison of KWATT to the original WATT and consider benefits/limitations of using KWATT for cases to which WATT was previously applied. This work is supported by NSF through the ITR grant NSF-0426867. [1] Virtual Laboratory for Earth and Planetary Materials, http://vlab.msi.umn.edu, September 2007 [2] C++/Tcl Library, http://cpptcl.sourceforge.net, September 2007 [3] TclSOAP Library, http://tclsoap.sourceforge.net/, September 2007
IN11B-0469
A System for Scripted Data Analysis at Remote Data Centers
Terascale data reduction and analysis still remain elusive for most geoscientists even as large scale geophysical models and observing systems begin to produce petascale datasets. Massive amounts of geophysical netCDF data remain underutilized due to scientists' limited bandwidth and computational capacity. Compounding the problem are the unique characteristics of geoscientists' data analyses which are poorly matched to traditional grid technologies. Our system, the Script Workflow Analysis for MultiProcessing (SWAMP) allows scientists to efficiently leverage grid resources for parallel computation through a rich scripting syntax compatible with most existing analysis scripts. Although shell script languages are impossible to optimize for all cases, data analysis scripting is well-defined sufficiently to allow inference of dependencies and use of advanced compilation techniques. SWAMP compiles scripts to detect potential parallelism and to dynamically schedule for parallel execution on resources allocated by an industry-standard grid scheduling engine. By its unique focus on I/O and bandwidth considerations rather than raw computational power, our system more efficiently handles the characteristic I/O-boundedness of geoscience workflows. Although scientists' scripts exhibit significant available parallelism, such parallelism is poorly exploited by conventional means that manage I/O and data locality as afterthoughts, if at all. In contrast, SWAMP eliminates unnecessary data movement to achieve its parallel efficiency in two ways: by providing a computational interface integrated with data service, and by scheduling computation to avoid I/O bottlenecks in disks and networks. This efficiency and performance is designed to integrate with scientists' analysis scripts-- many existing scripts require no modification to work with SWAMP. We test our system in uniprocessor, multiprocessor, and clustered environments, with both simple and terascale test cases. Benchmarks and other workflow statistics quantify not only the real performance benefits but also the reasons that render traditional methods unsuitable. We show that SWAMP typically modestly increases the data source's computational load while significantly reducing bandwidth in nearly all cases, making it suitable for installation in small lab groups as well as large data centers. Local computation and bandwidth requirements are thus drastically reduced, freeing the scientist to perform exploratory analysis and discovery in wider scopes and finer resolutions, and reducing time to discovery. http://code.google.com/p/swamp/