2018 Proposal Text
| Goal | Developing/adapting innovative technologies |
Utilizing developments to explore and understand ocean systems |
Technology/knowledge transfer |
|---|---|---|---|
| Percentage | 75% | 25% | 0% |
2b) Continuing Projects
Current status of project to date
In the 2017 proposal, we defined two main goals that were then broken down into smaller tasks. These goals are listed below with the work done in the first 6 months of 2017.
- Build the advisory capacity of the Data TAG (essentially the 'A' in TAG).
- We had discussions with BluHaptics regarding their data infrastructure and they are heavily focused on highly tailored data sets for use in machine learning. We definitely see their technology as a possible consumer for our data sets, but we did not see any specific technology that would help with the first steps of defining an architecture for storing, cataloging and accessing hetergeneous data sets.
- Held a teleconference with the OOI CI lead at Rutgers, Manish Prashar, and had a good discussion about requirements and how they met the requirements for the OOI CI. The core technology used (uFrame) was somethign that was utilized by the contractor (Raytheon) in the rush to take over development from the initial UCSD-led team. They made reference to uFrame being something that was not intended to be used by the outside community and, in fact, might be replaced with something later in time. It was a known quanity for getting the CI deployed, but was not intended for others to utilize as their own infrastructure.
- We had a few discussions with the CenCOOS data infrastructure team, Axiom, led by Rob Bocheneck, and while their infrastructure utilized many technologies that we have been researching (Highly parallel servers running software stacks like Elasticsearch, OPeNDAP, THREDDS, etc.) they were completely consumed by their commitments to the various IOOS projects and MBON (among others) and we could not get any of their time (we even tried to pay money for their time through the MBON project).
- We talked with attendees of the Signatures Discovery Institute and, similar to BluHaptics, they are focused on analysis and not the core infrastructure of data management. They were helpful with feedback from a client's perspective which will be helpful in future architecture work
- Attended the Renewable Natural Resources Foundation's Congress on Harnessing Big Data for the Environment. Really great talks and the cloud infrastructure and analysis stacks that Microsoft were actually quite impressive. Had a really good discussion with the presenter Kristin Tolle and she was interested in our work. Follow up requests have gone unanswered, but will keep up the research. Also got to hear from Jeff de La Beaujardière on how NOAA is dealing with their large data sets. While their requirements are different, their experience is valuable. With the size of NOAA they were able to procure very good relationships with cloud providers as they were intrested in hosting and serving NOAA's data as a way to attract data processors to run their analysis on their cloud infrastructure since the data is too large to move around. Conference report can be found here: http://www.rnrf.org/rrj/RRJV30N4.pdf.
- From the Congress on Harnessing Big Data meeting, we had a follow up teleconference with Shelley Stall form AGU about their Data Management and Assessment Program. Shelley was going to work on a scope of work for this assessment, but it looked to be in the $100K range which is not feasible for MBARI at this time.
- Attended Blue Tech Week in San Diego and while there were good discussions, outside of one specific day for "Big Data", it was largely focused on industrial businesses and venture capital. The day on Big Data was informative. Several talks on situational awareness and client side data analytics, but only a couple on data infrastructure. No real leads came from this meeting.
- Attended Elasticsearch training which seems like a very applicable technology and will be researched and prototyped in the future.
- Attended a HADOOP workshop and while the technology is interesting, it often requires a very large support infrastructure and may not be well aligned with our requirements. More research will be needed as time moves on to keep track of developments with HADOOP.
- Webinar on Lamba Architectures as a possible architecture for high bandwidth streaming data and real time analytics
- Webinar on SGCI's science gateways. This group could have valuable insight from a requirements perspective and we will be doing more research in this area.
- Early research on other technologies such as Google's Big Table and Map Reduce.
- Develop models, architectures, and technology recommendations and define a data technology roadmap for MBARI. In order to do this, we broke the work into several sections
- Define requirements and success criteria
- For this, several attempts have been made to gather information from what seems like our top priority stakeholders which were are the Packard Foundation and the Center for Ocean Solutions. To date, we have had no engagement from those resources and are trying to determine the next steps in gathering requirements from these stakeholders. We took the stakeholder work done by our external contractor as part of the External Web Upgrade Project as an inital discussion point for requirements planning. These stakeholders are:
- Foundation and Board Members
- Policy Makers and Resource Managers
- Researchers (External and Internal)
- Engineers
- Aquarium, Public, Students
- Teachers
- Media
- Tech Transfer Partners
- While we started the work of requirements gathering from existing systems (lessons learned) and future development prospects (new projects), we felt that our highest priority should be our top stakeholders which will be the focus of the next steps in the Data TAG work.
- For this, several attempts have been made to gather information from what seems like our top priority stakeholders which were are the Packard Foundation and the Center for Ocean Solutions. To date, we have had no engagement from those resources and are trying to determine the next steps in gathering requirements from these stakeholders. We took the stakeholder work done by our external contractor as part of the External Web Upgrade Project as an inital discussion point for requirements planning. These stakeholders are:
- Define requirements and success criteria
- Develop recommended pathways for data export to the outside world.
- We spent time researching data serving technologies and are using some of these technologies in projects to further understand their capabilities. Some of the technologies include ERDDAP, THREDDS, REST, OGC standards, etc.
- Define architecture designs for services, security, and large data set management that will enable the discovery and access requirements needed for the Technology Roadmap Challenges of data merging, integration, assimilation and data mining.
- As mentioned, above, we reached out through many different avenues to research architectures being used in computing today. There is still much work to do here and more will be done in the rest of 2017.
- Auth0 for chosen as our initial security provider to help with authentication.
- We are working with IS to understand how the eventing system on the Isilon cluster operates so that we can utilize the most efficient mechanism of cataloging data stores.
- Define recommended protocols and standards for data and descriptions
- We developed a REST Design Guide for service definition (https://docs.mbari.org/rest-design-guide/) for developers
- We developed a JSON Format Guide (https://docs.mbari.org/rest-design-guide/json/)
- We have chosen to use Markdown and MKDocs for documentation
- We have chosen Swagger for API Documentation
- Define recommendations for catalog and discovery technologies and architectures
- We are researching how other groups do their catalog systems and how our internal and external clients are trying to find an access our data. We have been working with catalog systems like ERDDAP, THREDDS, and Elasticsearch in our existing projects and along with existing internal systems to try to understand what concepts are critical in a catalog system.
- Define data models and recommended standards
- We have been doing research into standard groups to evaulate their data models against our needs. These groups include OGC, Unidata, OceanSITES, etc.
- We are starting to develop the highest level abstract models base on concepts like 'activities' and 'resources'
- Recommendation for policy and process and institute requirements
- Without the engagement from our highest priority stakeholders, this work has not moved forward.
Significant changes from previous project proposal
Through the first half of 2017, there has been an emerging message from the Packard Foundation, the Board and the Management Team searching for a way in chich we can communicate a measure of how MBARI's work (and the Foundation's investment) are impacting the world. The Data TAG team felt that this was an important direction that should have emphasis in the very near term. We are proposing to slow some of the architecture research and infrastructure development in exchange for a concerted effort between the Management Team, Accounting, ITD and Information Engineering, to focus on the requirements gathering effort for this new direction. We are proposing an interative prototype development effort between these groups to try and grasp how we can communicate our impact to our most important stakeholders, the Foundation and the Board. If we can create a tool to communicate this information effectively, that will serve as the foundation for future development to meet the needs of the rest of the stakeholders.
2c) New and Continuing Projects: Plans for 2018
Research Activities
The primary goal of the Data TAG has not changed from prior years. The Data TAG still holds the vision of providing an advisory and design capacity to make sure MBARI is tackling the 3 data-related challenges articulated in our Technology Roadmap:
- Developing methods for seamlessly merging molecular biological data with other environmental observations
- Developing methods for assimilating empirical observations into four-dimensional, coupled physical and biological models
- Developing tools for preserving, exploring, and mining multi-disciplinary data sets.
Figure 1 - Requirements Flow
Figure 1 shows a diagram of the flow of requirements gathering that has been followed to date for the Data TAG. During the External Web Upgrade project, we worked with an external design contractor to identify MBARI's key stakeholders. It was unanimous that for our external web presence, the highest priority stakeholders were our funding source and our board. We feel the same goes for the Data TAG efforts. Work to elicit requirements from key stakeholders has not been successful to date, we feel, due to a difficulty in articulating what it is we need to communicate for the board and the foundation to feel their investment has a positive impacts on their goals. Instead of the direct Q&A approach to requirements gathering from the board and the foundation, the Data TAG is proposing a different approach to try and get past this difficult part of the development process.
In order to effectively gather requirements about what meets our stakeholders needs, we feel an iterative, prototype driven approach is the best option to help extract requirements from a space that is not well defined. This is taking Chris and Mike's idea of 'bringing rocks' to a prototyped product that can be used to try and extract requirements.
Engineering/Technology Development
For this approach to work, we are proposing the following process.
- The Data TAG group will work to survey and compile features from as many different story-telling data portals as we can (maybe even outside the domain of oceanography). The group will sit down and compile thoughts about these different implementations (strengths/weaknesses) and then draw up what might be a simple first round prototype using MBARI data and information. As an example, Figure 2 shows a screen capture of the interactive funding map Chris showed in his Director's Dialog.
Figure 2 - Interactive Funding Map - After this initial stage of research and prototyping, we will start to work on 2-4 week cycles where we meet with the Management Team and ITD to demostrate and discuss features and possible new directions. The Data TAG team will then take this feedback and implmement another round of development to try out new ideas. The Data TAG group will keep track of requirements and features as this process continues so that after this process is complete, we will have both a set of concrete requirements and a prototype that demonstrates these features. The example show below in Figure 3 was a test application built for some catalog development work under a different project, but shows an example of what type of tools we are considering. The figure shows circles of where MBARI data collection activities have taken place. The different colors and size represent the numbers of activities that have taken place in that general area. This data was mined from the Expedition Database, STOQS, SSDS, Samples Database and BOG.
Figure 3 - Example of Prototype - During this iterative process, the Data TAG will also be developing models to capture the requirements and features needed for the prototype.
- At this point, we hope to then extend the iterations to include feedback from board members and the foundation members to expand those requirements to include those parts of the organization.
One thing to note is that we are planning on focusing on this iterative development process as the sole focus of the Data TAG for 2018 while some aspects of the work done to date (listed under 'Current status of project and results to date') will be moved to the CANON, ESP, and Integrated Time Series projects so that work can continue under projects that have more well defined requirements.
Deliverables
The deliverables for this work will be:
- Prototypes that test ideas and foster dicussions to target requirements gathering for an interface that can convey data-driven "stories" to our most important stakeholders
- Requirements gathered from our stakeholder (Foundation and Board) based on their experience with the prototype
- A roadmap that can be used to plan future work both on the interace and to plan work to stitch the interface work with the backend work that continues with other projects and data systems.
- Models that can capture information needed to support our development work.
Milestones
- Mid January: Compilation of features from other systems
- Early February: First prototype
- Every 2-4 weeks: Development/feedback iteration
- July: Requirements from iterative process to help defined future work
Tasks and Labor:
- Survey of other systems: 40 hours each engineer
- Each iteration: 1 week of each developer (6-7 iterations)
- Proposal review: 10 hours each
- Project Management: 40 hours me
Budget Justification
- Travel:
- 4 Site visits at $2000 each: $8000
- 2 conferences at $4000 each: $8000
- Cloud provisioning costs: $5000
Project Team
- Danelle Cline
- Duane Edgington
- Kevin Gomes
- Mike McCann
- Carlos Rueda
- Karen Salamy
- Brian Schlining
- Rich Schramm
- Management Team
- ITD
Data Storage Requirements:
- None: will be using cloud hosting
Data Availability
- The protype will be available to the Management Team, ITD and any Board and Foundation members we can get onboard.
Other info
Chris' feedback
- Really liked the idea.
- Wanted to mention to the board this November to prep them to be involved
- Focus: Trend Analysis (looking to the future)
- Add a question to the first two requirements: "What's the future?"
- Focus: What do we have to show for it?
- Focus: What was enabled by this work/investment?
- He was thinking that we might do it in terms of the three categories on the proposal (not sure how the heck we would do that).
- Maybe have 'regional' maps that are focused on telling stores.
- Really liked the idea of integration with ITD (new director)
- Really liked the survey, iterate approach
- Was thinking about whether he might work directly with Meg Caldwell and Dona Crawford to get them involved. Might visit the foundation this week.
- Emphasize the connection between PF Goals and what we do.
- Liked the idea of integrating data from other PF grantees as well.
- Everything should be space and time indexed, but could link to other non-spatial data
- REALLY liked the idea of having the 'story' up front (driven by ITD) and then driving down into details (i.e. products and data)
- Thought the idea of splitting catalog and this development was the right thing to do to keep focus.