Showing posts with label data preservation. Show all posts
Showing posts with label data preservation. Show all posts

Friday, February 27, 2015

Data on the Web Best Practices

Data on the Web Best Practices. W3C First Public Working Draft. 24 February 2015.
This document provides best practices related to the publication and usage of data on the Web. Data should be discoverable and understandable by humans and machines and the efforts of the data publisher recognized.This will help the interaction between the publishers and users.

Data on the Web allows for the existence of multiple ways to represent and to access data which is a challenge. Some of the other challenges include: metadata, formats, provenance, quality, access, versions, and preservation. The Best Practices proposed should help data publishers and data consumers overcome the different challenges faced during the data life cycle on the web. The draft proposes best practices for each one of the described challenges.

Tuesday, February 17, 2015

AHRQ Public Access to Federally Funded Research

AHRQ Public Access to Federally Funded Research. Francis D. Chesley. Agency for Healthcare Research and Quality. February, 2015.
The Agency for Healthcare Research and Quality's  has established a policy for public access to scientific publications and scientific data in digital format resulting from funding through the agency. Preservation is one of the Public Access Policy's primary objectives.

The Public Access Policy includes the following objectives:
  • Ensure that the public can access the final published digital documents.
  • Facilitate easy public search, analysis of and access to these publications
  • Ensure the attributes to authors, journals, and original publishers are maintained.
  • Ensure that publications and metadata are in an archival solution.
  • Ensure that all researchers receiving grants develop data management plans, describing how they will provide for long-term preservation of and access to scientific data in digital format.
Data management plans will include:
  • A plan for protecting confidentiality and personal privacy.
  • A description of how scientific data in digital format will be shared
  • It must include a plan for long-term preservation and access to the data
The data management plans will be evaluated based on the values of long-term preservation, access, and the associated cost, and administrative burden. AHRQ will contract with a commercial repository to ensure long-term preservation and full access to the public.

Digital scientific data is defined as "the digital recorded factual material commonly accepted in the scientific community as necessary to validate research findings including data sets used to support scholarly publications, but does not include laboratory notebooks, preliminary analyses, drafts of scientific papers, plans for future research, peer reviews, communications with colleagues, or physical objects, such as laboratory specimens."



Friday, January 23, 2015

The Dataverse Network

The Dataverse Network. Harvard Dataverse Network. 2014.
The Dataverse Network is an open source application to publish, share, reference, extract and analyze research data. It facilitates making data available to others and to replicate work of other researchers. The network hosts multiple studies or collections of studies, and each study contains cataloging information that describes the data plus the actual data and complementary files.

The Dataverse Network project develops software, protocols, and community connections for creating research data repositories that automate professional archival practices, guarantee long term preservation, and enable researchers to share, retain control of, and receive web visibility and formal academic citations for their data contributions.

Saturday, November 22, 2014

Five steps to decide what data to keep

Five steps to decide what data to keep. Angus Whyte. Digital Curation Centre. 31 October 2014.
 This guide aims to help UK Higher Education Institutions aid their researchers in making informed choices about what research data to keep. 

It will be relevant to researchers making decisions on a project-by-project basis, or formulating departmental guidelines. It assumes that decisions on particular datasets will normally be made by researchers with advice from the appropriate staff (e.g. academic liaison librarians) and taking into account any institutional policy on Research Data Management (RDM) and guidance available within their own domain.

Step 1. Identify purposes that the data could fulfill
Step 2. Identify data that must be kept
Step 3. Identify data that should be kept
Step 4. Weigh up the costs
Step 5. Complete the data appraisal 

The final step is to weigh the value of the data and any costs still to be incurred, "considering the long-terms aims, the qualities you identified, the time and money already invested in it and the risks of being unable to prepare any ‘must keep’ data for preservation."




Angus Whyte, Published: 31 October 2014
Angus Whyte, Published: 31 October 2014
Angus Whyte, Published: 31 October 2014
Angus Whyte, Published: 31 October 2014
Angus Whyte, Published: 31 October 2014

Saturday, March 29, 2014

Recommended Format Specifications

Recommended Format Specifications. Library of Congress. March 2014.
Recommended Format Specifications are hierarchies of the physical and technical characteristics of creative formats, both analog and digital, which will best meet the needs of all concerned, maximizing the chances for survival and continued accessibility of creative content well into the future.

There are two primary purposes of the specifications. One purpose of the specifications is to provide internal guidance within the Library to help inform acquisitions of collections materials (other than materials received through the Copyright Office). A second purpose is to inform the creative and library communities on best practices for ensuring the preservation of, and long-term access to, the creative output of the nation and the world.
Six broad categories of creative output, and particular format specifications in descending order of preference.
  • Textual Works and Musical Compositions
  • Still Image Works
  • Audio Works
  • Moving Image Works
  • Software and Electronic Gaming and Learning
  • Datasets/Databases
Format Specifications: PDF

Tuesday, June 4, 2013

Cerf sees a problem: Today's digital data could be gone tomorrow.

Cerf sees a problem: Today's digital data could be gone tomorrow. Patrick Thibodeau. Computerworld. June 4, 2013.
Vinton Cerf is concerned that much of the data that has been created in the past few decades and for years still to come, will be lost to time. Digital materials from today, such as spreadsheets, documents, presentations as well as mountains of scientific data, won't be readable in the years and centuries ahead. Software backward compatibility is very hard to preserve over very long periods of time, and the data objects are only meaningful if the software programs are available to interpret them. "The scientific community collects large amounts of data from simulations and instrument readings. But unless the metadata survives, which will tell under what conditions the data was collected, how the instruments were calibrated, and the correct interpretation of units, the information may be lost. If you don't preserve all the extra metadata, you won't know what the data means. So years from now, when you have a new theory, you won't be able to go back and look at the older data."

What is needed is a "digital vellum," a digital medium that is as durable and long-lasting as the material that has successfully preserved written content for more than 1,000 years. If a company goes out of business and there is no provision for its software to become accessible to others, all the products running that software may become inaccessible. The cloud computing environment may help; it may be able to emulate older hardware on which we can run operating systems and applications. We need to preserve the bits, but also the a way of interpreting them.

The CODATA Mission: Preserving Scientific Data for the Future

The CODATA Mission: Preserving Scientific Data for the Future.Jeanne Kramer-Smyth. Spellbound Blog. February, 2013.
This is a post (and a link to the slides) about a session that was part of The Memory of the World in the Digital Age: Digitization and Preservation conference. The aim was to describe the initiatives of the Data at Risk Task Group (DARTG), which is part of the International Council for Science Committee on Data for Science and Technology (CODATA).

The goal is to preserve scientific data that is in danger of loss because they are not in modern electronic formats, or have particularly short shelf-life. The task group is seeking out sources of such data worldwide since many are irreplaceable for research into the long-term trends that occur in the natural world. One speaker talked about two forms of knowledge that we are concerned with here: the memory of the world and the forgettery of the world. Only the digital, or recently digitized, data can be recalled readily and made immediately accessible for research in the digital formats that research needs. The “forgettery of the world” is the analog records, ones that have been set aside for whatever reason, or put away for a long time and have become almost forgotten.  It the analog data which are considered to be “at risk” and which are the task group’s immediate concern.  Some of the early digital data are insufficiently described, or the format is out of date and unreadable, or the records cannot be located at all easily.

How can such “data at risk” be recovered and made useable?  An inventory website has been set up where one can report data-at-risk. The overarching goal is to build a research knowledge base that offers a complimentary combination of past, present and future records. Some data mentioned: Oceanographic; climate; satellite; and other scientific data sets; born digital maps. With digital preservation initiatives there is a lot of rhetoric, but not so much action. There have been many consultations, studies, reports and initiatives but not very much has translated into action. 


Monday, May 6, 2013

The APTrust Architecture Presentation.

The APTrust Architecture Presentation. Scott Turnbull. Academic Preservation Trust. May 6, 2013.
The Academic Preservation Trust (APTrust) consortium is developing a preservation environment.  The website includes slides presenting the APTrust Phase I Architecture.  It gives a general look at the components being developed.  The APTrust repository will serve as a replicating node for the Digital Preservation Network (DPN). At the local level, APTrust will provide a preservation environment for participating members, including disaster recovery services.



Saturday, May 4, 2013

DuraCloud Now Offers Low Cost Glacier Storage!

DuraCloud Now Offers Low Cost Glacier Storage!  Carol Minton Morris. Duraspace. May 2, 2013.
DuraCloud has increased the long-term storage options that it offers.  This online storage is intended for durable storage, for data archiving and backup, and particularly for data that is infrequently accessed. It has automatic synchronization between primary and secondary copy, and web access to all copies stored in DuraCloud. Pricing is available at http://duracloud.org/pricing.


Monday, April 22, 2013

Stakeholder Benefits from Research Data Management: new document from Research360 project.

Stakeholder Benefits from Research Data Management: new document from Research360 project. Neil Beagrie, Catherine Pink. University of Bath. 
The Research360 Project has released the summary stakeholder benefits analysis from the Research Data Management business case for the University of Bath. The 4 page document is available for  download in PDF format.

Industry and private sector partnerships alongside public sector and voluntary sector partnerships are key elements of many university research programmes.  Frequently partners sharing their practice, results data and laboratory methodologies can lead to vital knowledge transfer activities, improved services and products, creation of spin-out companies and further investment in the Higher Education sector.  A summary list of stakeholder benefits that can arise from research data management in these collaborations. Benefits are listed for:
  • university community by its key stakeholder groups:
    • academic staff and researchers, students, professional services, and the institution
  • external partners:
    • industry and commerce, public/voluntary sectors, government, and society 
 Some of the benefits include:
  • Improve possibility of success in research funding by addressing any concerns around data management.
  • Safeguarding your data against potential loss.
  • Support in patent issues such as proof of provenance through improved use of version control.
  • Enhanced global reputation through recognition of the quality of research outputs and data infrastructure.
  • Attract new collaborators and accelerate deepening of existing relationships.
  • Graduate employability increased through university partner connections and student data skills.
  • Reduction of risk for sensitive data if data transfer is secure.
  • Cost efficiencies from shared data services.


Wednesday, March 27, 2013

Supporting the Changing Research Practices of Chemists.

Supporting the Changing Research Practices of Chemists.  February 25, 2013. Matthew P. Long, Roger C. Schonfeld. Ithaka S+R. February 26, 2013. [PDF]

This report, intended for those who support chemists, including librarians, is about the latest research methods, practices, and information services needs of academics chemists. Chemists need services to make their lives easier and their research groups more productive; this includes minimizing paperwork and administrative tasks. They value academic libraries primarily for the access that they provide to electronic journals and other online resources. Researchers are often frustrated by an inability to share large amounts of data with a collaborator. Few chemists visit the physical library, but they use the library digital collections heavily.

In the survey, fewer than 10% reported a research consultation with a librarian, asked for help with a data management, or asked for assistance on an issue related to publishing in the past year; they rarely reach out to the library to discuss issues or request support. The main search sites for chemists are Web of Knowledge/Web of  Science, SciFinder, and PubMed. It would be helpful to have tools to help process all of this information,  a pre-scan of announcements from journals, and organize their materials. Electronic Lab Notebooks (ELNs) make it easy to share, archive, and search through past lab notes, but are at risk in the lab. Labs generally do not have good data management infrastructure or proper external support for developing it, especially in sharing and preserving files.

It is difficult for academic chemists to coordinate the recording and preservation of data after the completion of a project. When data are saved, they are often held in unstable or at-risk formats  or in formats where no one else can access or interpret them. Sometimes a large amount of potentially useful data is not shared or preserved in any durable way. One chemist invited the library to come and speak to the department about preservation and access. Chemists have a general lack of awareness of  effective data curation and preservation. Data management and preservation is time-consuming and rarely straightforward; it requires expert advice and constant monitoring.

The findings:
  1. Chemists need better support in data management, sharing and preservation.
  2.  Many researchers remain anxious about keeping up with the newest literature.
  3. They need new tools to stay aware of new research and also serendipitous discovery.
  4. Chemists  require greater support in disseminating their research, including articles, data, and other materials.
Other areas of concern for academic chemists : laboratory management, gaining access to industrial funding, and teaching support.
We see some real potential for the academic library to stretch the definition of the services it offers to the academic chemist. The library may also have a role in working with other service providers and ensuring that academics are aware of the latest research tools. It is clear from this project that libraries must think strategically about whether and how to invest in services for chemists.
 

Friday, November 16, 2012

The Data Conservancy Instance: Infrastructure and Organizational Services for Research Data Curation

The Data Conservancy Instance: Infrastructure and Organizational Services for Research Data Curation. Matthew S. Mayernik, et al. D-Lib Magazine. September/October 2012.
Digital research data can only be managed and preserved over time through a sustained institutional commitment. Digital research data, if curated and made broadly available, promise to enable researchers to ask new kinds of questions and use new kinds of analytical methods in the study of critical scientific and societal issues. 

The Data Conservancy, a community organized around data curation research, technology development, and community building, is driven by a common theme: the need for institutional solutions to digital research data collection, curation and preservation challenges.
The four main activities of the Data Conservancy are:
  1. A focused research program to examine research practices across multiple disciplines in order to understand the data curation tools and services needed to support interdisciplinary research
  2. An infrastructure development program for data management and curation services
  3. Data curation educational and professional development programs 
  4. Development of sustainability models for long term data curation.
Data curation solutions for research institutions must address both technical and organizational challenges. These include context, hardware and software infrastructure, services, and sustainable strategy.

The features needed include a preservation-ready system, customizable user interfaces, flexible data model, ingest and search interface,  and data examination processes.

Friday, September 7, 2012

Data Management Plans

Data Management Plans. Website. MIT. 2012.
Many grants and programs are requiring a data management plan. This site gives some great advice on how to create a plan and to manage your data. "A data management plan will help you to properly manage your data for own use, not only to meet a funder requirement or enable data sharing in the future." The plan  components may include:
  • Description and purpose of the project
  • Description of the data and how it will be collected
  • Standards to be applied, especially the file formats and the metadata
  • Plans for the short term data management and the long term data archiving
  • Data access policies
  • Responsibilities of, and people involved in, the data management
The site also contains checklists and other tools for data management.


Monday, August 13, 2012

The Problem of Data

The Problem of Data. Lori Jahnke, Andrew Asherpub, Spencer D. C. Keralis. CLIR Report. Council on Library and Information Resources. August 12, 2012.
Excellent report on data storage, use, and curation.  A section contains a snapshot of the current digital data curation education landscape.  Below are some long notes and excerpts from the PDF article:

Key Findings
  • None of the researchers interviewed for this study have received formal training in data management practices, nor do they expresssatisfaction with their level of expertise.
  • Few researchers, especially among those who are early in their career, think about long-term preservation of their data.
  • The demands of publication output overwhelm long-term considerations of data curation. Metadata and documentation are of interest only if they help a researcher complete his or her work.
  • There is a great need for more effective collaboration tools, as well as online spaces that support the volume of data generated and provide appropriate privacy and access controls.
  • Few researchers are aware of the data services that the library might be able to provide and seem to regard the library as a dispensary of goods (e.g., books, articles) rather than a place for research/professional support.
Recommendations
  • There is unlikely to be a single out-of-the-box solution that can be applied to the problem of data curation. Instead, an approach is needed that emphasizes working with researchers to identify or build appropriate tools.
  • Researchers must have access to adequate networked storage.
  • Universities should revise access policies to support multi - institutional research projects.
  • Programs should begin early in the researcher career path for the greatest long-term benefit.
  • Data curation systems should be integrated with the active research phase (i.e., as a backup, etc).
  • Privacy and data access control tools should be developed to manage confidential data. Policies must be developed that support researchers in using these technologies.
Other notes:
  • Data curation, a term generally defined as a set of activities that includes the preserving, maintaining, archiving, and depositing of data to keep it secure, intact, and accessible for reuse.
  • Many researchers expressed concerns surrounding the ethical reuse of research data. Additional work is needed to establish best practices in this area, particularly for qualitative data sets.
  • Most participants reported feeling adrift when establishing protocols for managing their data and added that they lacked the resources to determine best practices, let alone to implement them. Almost none of the scholars reported that data curation training was part of their graduate curriculum.
  • Perhaps one of the more complicated issues for data curation is the complex life cycle of research data and projects. Data collection may occur throughout the project and change from before it is completed.
  • Scholars may collect data on a phenomenon unrelated to their current project with no clear idea of the potential usefulness of those data. Such data might be integrated with a later project, given away to an interested colleague, or never used at all.
  • It would be helpful to have a way to collect data into a collection space that could be used throughout the project.
  • The researchers held contradictory views about the value of their data. Some wanted to associate their data with publications or to have it available for use in the classroom
  • Few of the researchers thought about long-term preservation of their data, especially those who were early in their career.
  • The academic system offers little or no career reward for preserving one’s data.
  • Data preservation strategies must take into account varied, proprietary, and non-standard data formats, and provide a real-time benefit for the scholar in meeting research goals.
  • Given the lack of infrastructure for sharing and storing data, the social sciences may face similar problems of data loss in documenting social phenomena as researchers begin to work within larger collaborative groups and with larger data sets. Data stored on personal media devices are especially vulnerable to this type of loss, as few scholars have the skills necessary to maintain data over time and across hardware and software platforms. Several of the scholars interviewed reported storing data on legacy systems that may become inaccessible
  • University policies that appropriately address the ethical considerations relating to data sharing and preservation would benefit researchers, administrators, and technologists alike.
  • Researchers hold tremendous amounts of data on personal computers and hard drives, many of which are not backed up adequately. Among the participants, the research data ranged from under 1 GB to multiple terabytes. Data types included various formats of images, video, audio files, data sets, documents, etc.
  • Managing large files presents significant challenges for researchers in that university infrastructures typically do not provide adequate storage space or sufficient bandwidth for data access.  The data may be lost when researchers upgrade their computers or software. Few researchers put more than minimal effort into organizing non-active data or ensuring its continued compatibility with new software or hardware.
  • There is a clear need for libraries to move beyond passively providing technology to embrace the changes in scholarly production that emerging technologies have brought.  
  • The data preservation step must be fully integrated into a scholar’s research workflow. Not only are necessary metadata and other materials much more easily captured while research is in progress, but also there is a real opportunity to streamline research workflows and to provide much needed support. Scholars need help with the technical aspects of managing and preserving data, as well as with basic curation issues (e.g., what to keep and what to delete), and the ethical implications of sharing their data (e.g., what is an appropriate latency period for the data and how does one balance the need to provide meaningful access with the risk of inadvertently exposing confidential participant information).
  • Although some researchers acknowledge that their data could be useful to other researchers, there is little incentive to invest time in archiving or repackaging data sets.
  • Extensive outreach to scholars is necessary to build the relationships that will facilitate data preservation. This is likely to be a slow process initially. Researchers are unlikely to engage with those they do not view as peers.
  • Researchers need additional tools to manage preserved data on their own, and they would benefit from access to professionals who can offer advice on management strategies.
  • Researchers typically align themselves with their disciplines rather than with their institutions; therefore, support models that extend beyond the university are likely to be especially beneficial.
  • Reaching the level of collaboration among universities and the technical interoperability required to capture and preserve a career’s worth of data in the current environment is a challenge.
  • Current data management systems must be fundamentally improved so that they can meet the capacity demand for secure storage and transmission of research data. Integrating the data preservation system with the active research cycle is essential to encourage researcher investment.
  • Researchers are not well positioned to meet the technical and policy challenges without the coordinated support of libraries, information technology units, and professionals who possess both technical and research expertise.
  •  One example concerning the PETRA e+e collider project in Hamburg, Germany; In the more than 25 years since, theoretical insights and computing advancements have made the data valuable once again. However, much of the data have been irrevocably lost to corrupt storage media, lost computer code, and deactivated personal accounts. These early particle physics experiments are unique, as modern colliders operate at higher energy levels and cannot replicate the particle interactions.

Wednesday, May 16, 2012

Implementing DOIs for Research Data.

Implementing DOIs for Research Data.  Natasha Simons. D-Lib Magazine. May/June 2012.
As research becomes more collaborative and global it is also becoming more difficult to manage the large amounts of research data generated daily.  The Digital Object Identifier (DOI) system is one way to create persistent identifiers for research data collections and datasets. "Data that is richly described, organised, integrated and connected allows the data to be more easily discovered by other researchers." Identifying such resources allow research data collections and datasets be open and discoverable to others, but there are questions that need to be answered, such as the type of material to get a persistent id, the granularity, whether the landing page or the resource should get the id, who creates and maintains the ids, and for how long. The questions, common to other institutions, should encourage discussion and collaboration.

Thursday, December 8, 2011

Why don't we already have an Integrated Framework for the Publication and Preservation of all Data Products?

Why don't we already have an Integrated Framework for the Publication and Preservation of all Data Products?   Alberto Accomazzi,et al. Astronomical Data Analysis Software and Systems.  
7 Dec 2011.
Astronomy has long had a working network of archives supporting the curation of publications and data. There are examples of websites giving access to data sets, but they are sometimes short lived.  "We can only realistically take implicit promises of long-term data archival as what they are: well-intentioned plans which are contingent on a number of factors, some of which are out of our control." We should take steps to ensure that our system of archiving, sharing and linking resources is as resilient as it can be.  Some ideas are: 
  1. future-proof the naming system: assign persistent data IDs to items we want to preserve 
  2. provide the ability to cite complete datasets, just as we can cite websites
  3. include a data reference section in academic papers
Curated datasets need to be preserved indefinitely for scholarly purposes.

Friday, September 16, 2011

Long-term Preservation for Spatial Data Infrastructures: a Metadata Framework and Geo-portal Implementation.

Long-term Preservation for Spatial Data Infrastructures: a Metadata Framework and Geo-portal Implementation. Arif Shaon, Andrew Woolf. D-Lib Magazine. September/October 2011.
Geospatial data is increasing, particularly with diverse environmental datasets. Long-term preservation of the data is not typically addressed, but it is very important for current and future use.  Sustained access to environmental data is becoming more important and more difficult because it is increasing so dramatically.
Without effective long-term preservation, the data face the risk of becoming unusable over time. This article looks at the requirements, particularly metadata, for preserving this data.  The authors have implemented a web-based portal prototype that demonstrates some functions of a preservation interface, such as data discovery using geospatial metadata, data downloading, metadata creation and validation.  There is more to be done in this area.

Sunday, September 4, 2011

JISC Legal Cloud Computing and the Law Toolkit.

JISC Legal Cloud Computing and the Law Toolkit. Website. 31 August 2011.
Documents to help make informed decisions about implementing cloud computing solutions in an  institution. Not specifically for digital preservation, but helpful to think about the policies that will affect data and the life-cycle, such as
Data Protection, Possession of Data on Termination, etc.  Written for UK educational institutions.

  • Report on Cloud Computing and the Law for UK Further and Higher Education

  • User Guide: Cloud Computing and the Law for IT

  • User Guide: Cloud Computing and the Law for Senior Management and Policy Makers

  • User Guide: Cloud Computing and the Law for Users 

  • User Guide: Cloud Computing Contracts, SLAs and Terms & Conditions of Use 

Institutional Repository and ETD Bibliography 2011

Institutional Repository and ETD Bibliography 2011. Charles W. Bailey, Jr.  September 2011.
This bibliography has over 600 English-language articles, books, and other works about institutional repositories and theses and dissertations (ETDs).  Among other things, it includes digital preservation issues, IR library issues, IR metadata strategies, and institutional open access mandates and policies. Most sources have been published from 2000 through June 30, 2011.  The bibliography includes links to freely available versions of included works.  It is available as a PDF file.

Monday, August 29, 2011

Criteria for the Trustworthiness of Data Centres

Criteria for the Trustworthiness of Data Centres. Jens Klump. D-Lib Magazine. January/February 2011.
The rapid decay of URLs for research resources is an important reason to use persistent identifiers. The use of persistent identifiers implies that the data objects are persistent themselves. The rapid obsolescence of the technology to read the information, along with the physical decay of the media, represents a serious threat to preservation of the content. Since research projects only run for a relatively short time, it is advisable to shift the responsibility for long-term data curation from the individual researcher to a trusted data repository or archive.

We need criteria for the assessment of trustworthiness of digital archives. Some of the methods presented have been:
  •     Trustworthy Repositories Audit & Certification: Criteria and Checklist (TRAC)

  •     Catalogue of Criteria for Trusted Digital Repositories (nestor Catalogue)

  •     DCC and DPE Digital Repository Audit Method Based on Risk Assessment (DRAMBORA)

  •     DINI-Certificate Document and Publication Services

  •     Data Seal of Approval (Sesink et al., 2008)

These provide useful feedback on developing additional criteria and auditing procedures to certify  trusted digital archives.