Showing posts with label curation. Show all posts
Showing posts with label curation. Show all posts

Tuesday, March 10, 2015

Investing in Curation. A Shared Path to Sustainability. Final RoadMap.

Investing in Curation. A Shared Path to Sustainability. Paul Stokes. The 4C project. March 9, 2015.
Digital curation involves managing, preserving and adding value to digital assets over their entire life cycle. Actively managing digital assets maximizes their value and reduces the risk of obsolescence. The costs of curation is a concern to stakeholders. The final version of the road map is now available; it starts with a focus on the costs of digital curation, but the ultimate goal is to change the way that all organizations manage their digital assets.

The vision: Cost modeling will be a part of the planning and management activities of all digital repositories.
  • Identify the value of digital assets and make choices
    • Value is an indirect economic determinant on the cost of curating an asset. The perception of value will affect the methods chosen and how much investment is required.
    • Content owners should have clear policies regarding the scope of their collections, the type of assets sought, the preferred file formats.
    • Establish value criteria for assets as a component of curation, understanding that certain types of assets can be re-generated or re-captured relatively easily, thereby avoiding curation costs
  • Demand and choose more efficient systems
    • Requirements for curation services should be specified according to accepted standards and best practices.
    • More knowledgeable customers demanding better specified and standard functionality means that products can mature more quickly.
  • Develop scalable services and infrastructure
    • Organizations should aim to work smarter and be able to demonstrate the impact of their investments.
  • Design digital curation as a sustainable service
    • Effective digital curation requires active management throughout the whole lifecycle of a digital object.
    • Curation should be undertaken with a stated purpose.
    • Making curation a service further embeds the activity into the organization's normal business function.
  • Make funding dependent on costing digital assets across the whole lifecycle
    • Digital curation activity requires a flow of sufficient resources for the activity to proceed.
    • Some digital assets may need to be preserved in perpetuity but others will have a much more predictable and shorter life-span.
    • All stakeholders involved at any point in the curation lifecycle will need to understand their fiscal responsibilities for managing and curating the asset until such time that the asset is transferred to another steward in the lifecycle chain.
  • Be collaborative and transparent to drive down costs
    • Each organization is looking to realize a return on their investment.
    • If those who provide digital curation services can be descriptive about their products and transparent about their pricing structures, this will enhance possible comparisons, drive competitiveness and lead the market to maturity.

Monday, January 26, 2015

Digital Curation Foundations

Digital Curation Foundations. Stephen Abrams. California Digital Library. January 20, 2015. (PDF).
Digital curation is a complex of actors, policies, practices, and technologies that enables meaningful consumer engagement with authentic content of interest across space and time. The UC Curation Center defines its mission in terms of digital curation, rather than digital preservation because that better expresses the need for coordinated activities of preservation of, and access to, managed assets. It also reflects the idea of ongoing enrichment of managed content rather than just maintaining the content over time. This should ideally start before the assets are created. The approach is more on services than systems, and those services should be delivered at the place they are needed.


The laws of library science deal with use, service to the users, and ongoing change. Every asset should be curated in order to be used. They should be able to be used when and where the user needs them and in accordance with the user's expectations. The digital curation activities must not only be sustainable, but capable of evolving to meet every changing needs as well as risks. This kind of service requires "administrative, financial, and professional support.

Curation decisions should be made with respect to an underlying theory or conceptual domain model based on first principles. The ultimate goal of digital curation is to deliver content. The digital curation field has reached a stage of maturity where it can usefully draw upon a rich body of theoretical research and practical experience. 

  • The curation imperative: providing highly available, responsive, comprehensive, and sustainable services for access to, and use and enhancement of, authentic digital assets over time.
  • The primary unit of curation management is the digital object
  • The true focus of curation is the underlying information meaning of the objects. "In other words, bits are the means, content is the ends."
Curation services include:
  • Creation / acquisition.
  • Appraisal / selection.
  • Preservation planning.
  • Preservation intervention.
  • Selection of appropriate curation service providers
  • Appropriate micro services

Monday, November 24, 2014

Curation Costs Exchange: Supporting Smarter Investments in Digital Curation

Curation Costs Exchange: Supporting Smarter Investments in Digital Curation. Sarah Middleton. Educause Review Online. November 10, 2014.

Tools to manage and estimate costs have not been integrated into other digital curation processes or tools. To determine why that is so a consortium of 13 European cost modeling specialists launched the Collaboration to Clarify the Costs of Curation (4C) project.

4C seeks to help organizations better understand the costs and benefits of digital curation and preservation, and to help users draw together existing and useful resources so they can both make their own assessment of existing models and develop their own cost modeling exercises. The Curation Costs Exchange (CCEx), a platform for the exchange and comparison of digital curation costs and cost information, is a key 4C project deliverable developed to support these goals.

The Cost Comparison Tool enables the exchange of sensitive data and gives users the opportunity to identify greater efficiencies, better practices, and valuable information exchanges among peers. There is also the Understand Your Costs toolkit. The Economic Sustainability Reference model highlights key digital curation concepts, relationships, and decision points in a complex problem space, helping users benchmark and compare their own local models.

Saturday, November 22, 2014

Five steps to decide what data to keep

Five steps to decide what data to keep. Angus Whyte. Digital Curation Centre. 31 October 2014.
 This guide aims to help UK Higher Education Institutions aid their researchers in making informed choices about what research data to keep. 

It will be relevant to researchers making decisions on a project-by-project basis, or formulating departmental guidelines. It assumes that decisions on particular datasets will normally be made by researchers with advice from the appropriate staff (e.g. academic liaison librarians) and taking into account any institutional policy on Research Data Management (RDM) and guidance available within their own domain.

Step 1. Identify purposes that the data could fulfill
Step 2. Identify data that must be kept
Step 3. Identify data that should be kept
Step 4. Weigh up the costs
Step 5. Complete the data appraisal 

The final step is to weigh the value of the data and any costs still to be incurred, "considering the long-terms aims, the qualities you identified, the time and money already invested in it and the risks of being unable to prepare any ‘must keep’ data for preservation."




Angus Whyte, Published: 31 October 2014
Angus Whyte, Published: 31 October 2014
Angus Whyte, Published: 31 October 2014
Angus Whyte, Published: 31 October 2014
Angus Whyte, Published: 31 October 2014

Thursday, October 30, 2014

Investing in Curation: A shared path to sustainability.

Investing in Curation: A shared path to sustainability. 4C Project. October 20, 2014.

Digital curation involves managing, preserving and adding value to digital assets over their entire lifecycle. The active management of digital assets maximises their reuse potential, mitigates the risk of obsolescence and reduces the likelihood that their long-term value will diminish. However, this requires effort so there are costs associated with this activity. As the range of organisations responsible for managing and providing access to digital assets over time continues to increase, the cost of digital curation has become a significant concern for a wider range of stakeholders.

Establishing how much investment an organisation should make in its curation activities is a difficult question. If a shared path can be agreed that allows the costs and benefits of digital curation to be collectively assessed, shared and understood, a wider range of stakeholders will be able to make more efficient investments throughout the lifecycle of the digital assets in their care. With a shared vision, it will be easier to assign roles and responsibilities to maximise the return on the investment of digital curation and to clarify questions about the supply and demand of curation services. This will foster a healthier and more effective marketplace for services and solutions and will provide a more robust foundation for tackling future grand challenges.

Situating the Roadmap:  The six messages in the roadmap have been carefully considered to effect a step change in attitudes over the next five years. It starts with a focus on the costs of digital curation—but the end point and the goal is to bring about a change in the way that all organisations think about and sustainably manage their digital assets.

D5.1 - Draft Roadmap ( PDF - 2.5 MB)


Saturday, March 23, 2013

Digital Curation Bibliography

Digital Curation Bibliography: Preservation and Stewardship of Scholarly Works, 2012 Supplement. Charles W. Bailey, Jr. March 2013.

Bibliography: Preservation and Stewardship of Scholarly Works, 2012 Supplement, which presents over 130 English-language articles, books, and technical reports published in 2012 about digital curation and preservation, copyright issues, digital formats (e.g., media, e-journals, research data), metadata, models and policies, national and international efforts, projects and institutional implementations, research studies, services, strategies, and digital repository concerns.
It is a supplement to the Digital Curation Bibliography:Preservation and Stewardship of Scholarly Works, which covers over 650 works published from 2000 through 2011.
 

Monday, February 11, 2013

Sustaining Our Digital Future Institutional Strategies for Digital Content

Sustaining Our Digital Future: Institutional Strategies for Digital Content. Nancy L Maron, Jason Yun and Sarah Pickle. Strategic Content Alliance. January 29, 2013. (PDF, 91 pp.)
The shift with digital media in scholarly communications is transformative; data sets, dynamic digital resources, websites, digital collections,  crowd sourced or born digital content: there are challenges and opportunities, along with questions about who is responsible for maintaining them, and how to maximize the value of the content.

Some findings:
  • project have received support from the host institution, but few have plans for ongoing support.
  • There are potential partners on campus, but project leaders do not seek them early
    enough when critical decisions are being made.
  • Digital projects across campuses may be hosted by many groups, which poses challenges for discovery. There is often no single place for users to find digital projects and some projects can too easily slip from view.
  • Current funding styles do not support ongoing operation
  • „Campus-wide solutions are beginning to emerge, but even these tend to address just the basic “maintenance” issues of storage, preservation and access.
  • „Focus is often on creating new content, with little thought about ongoing efforts to enhance the content or update user interfaces.
Recommendations:
  •  Perform an early and honest appraisal to find which projects are likely to require support after completion:
    1. Digital content requiring just “maintenance”: plan that the content will be deposited and integrated into some other site, database, or repository.
    2. „Digital resources requiring ongoing growth and investment: These require early sustainability planning, including identifying institutional or other partners and careful consideration of the full range of costs and activities needed to keep the resource vibrant.
  •  Be realistic in assessing the future needs of the resource at its outset and in continuing support.
  •  „Identify campus partners early on.
  •  „Consider how central your project is to the overall mission of the institution.
  • „Consider if projects could be drawn together to create a deeper network of support, both for “maintenance” projects and those with the potential to really grow.
  • Develop ways to help users find decentralized content and to reach out to content users. These could start as an inventory of all of the digital holdings or common catalogs.
  • „Determine where scale solutions pay off, where experts are best placed to champion a project, or create common storage, usage and preservation systems for an organization.
  • Continue to identify and support ongoing development of the “front-end”, including user needs, interface development, and content enhancement. Pay attention to the changing needs of users and determine what enhancements the digital resource will require.
Libraries, museums, technology departments, and digital humanities centres are among the players that have begun to emerge as potential leaders of greater coordinated digital support on university campuses. Libraries have begun to consider the support of digital resources to be a critical part of their missions. Some universities provide advisory support to project leaders and libraries provide help for digital projects, from help to understand grant requirements, co-developing projects, providing hosting, curation, and preservation expertise.

 Sustainability and Use: 
  •  Research data platforms: At some institutions major initiatives are underway to develop research data platforms. The goal of the platform is not just preservation and storage but access and reuse. The first step is to have a platform. From there they can test and refine the service for researchers depositing data sets, library  curated collections, and university departments.
  • "A coherent digital policy from early review, guidelines on costings and deposit standards, to forecasting what ongoing activities will be needed and who will carry them out, would ideally remove much of the risk of “digital time bombs” while obliging both project leaders and university leaders to take a moment to envision the ongoing impact they want these resources to have, and how to best achieve that."
  • Unlike universities (who often play the role of reluctant, passive, or simply unaware host, to a great deal of digital content created by their scholars) museums and libraries tend to be the ones initiating this work and are eager to build and maintain these collections.
  • Despite the benefits of centralisation, the mere presence of a catalogue and centralised
    repository does not ensure greater usage of or engagement with its holdings.
  • Many institutions devote considerable attention to the upfront creation of content, but not nearly as much to its ongoing enhancement or reuse, resulting in collections that are certainly present in the main catalogue, but otherwise exist only as capsules of content, frozen in time.
  •  Once the project is finished, management of the digital resource is not always clear.

Friday, November 16, 2012

The Data Conservancy Instance: Infrastructure and Organizational Services for Research Data Curation

The Data Conservancy Instance: Infrastructure and Organizational Services for Research Data Curation. Matthew S. Mayernik, et al. D-Lib Magazine. September/October 2012.
Digital research data can only be managed and preserved over time through a sustained institutional commitment. Digital research data, if curated and made broadly available, promise to enable researchers to ask new kinds of questions and use new kinds of analytical methods in the study of critical scientific and societal issues. 

The Data Conservancy, a community organized around data curation research, technology development, and community building, is driven by a common theme: the need for institutional solutions to digital research data collection, curation and preservation challenges.
The four main activities of the Data Conservancy are:
  1. A focused research program to examine research practices across multiple disciplines in order to understand the data curation tools and services needed to support interdisciplinary research
  2. An infrastructure development program for data management and curation services
  3. Data curation educational and professional development programs 
  4. Development of sustainability models for long term data curation.
Data curation solutions for research institutions must address both technical and organizational challenges. These include context, hardware and software infrastructure, services, and sustainable strategy.

The features needed include a preservation-ready system, customizable user interfaces, flexible data model, ingest and search interface,  and data examination processes.

Monday, August 13, 2012

The Problem of Data

The Problem of Data. Lori Jahnke, Andrew Asherpub, Spencer D. C. Keralis. CLIR Report. Council on Library and Information Resources. August 12, 2012.
Excellent report on data storage, use, and curation.  A section contains a snapshot of the current digital data curation education landscape.  Below are some long notes and excerpts from the PDF article:

Key Findings
  • None of the researchers interviewed for this study have received formal training in data management practices, nor do they expresssatisfaction with their level of expertise.
  • Few researchers, especially among those who are early in their career, think about long-term preservation of their data.
  • The demands of publication output overwhelm long-term considerations of data curation. Metadata and documentation are of interest only if they help a researcher complete his or her work.
  • There is a great need for more effective collaboration tools, as well as online spaces that support the volume of data generated and provide appropriate privacy and access controls.
  • Few researchers are aware of the data services that the library might be able to provide and seem to regard the library as a dispensary of goods (e.g., books, articles) rather than a place for research/professional support.
Recommendations
  • There is unlikely to be a single out-of-the-box solution that can be applied to the problem of data curation. Instead, an approach is needed that emphasizes working with researchers to identify or build appropriate tools.
  • Researchers must have access to adequate networked storage.
  • Universities should revise access policies to support multi - institutional research projects.
  • Programs should begin early in the researcher career path for the greatest long-term benefit.
  • Data curation systems should be integrated with the active research phase (i.e., as a backup, etc).
  • Privacy and data access control tools should be developed to manage confidential data. Policies must be developed that support researchers in using these technologies.
Other notes:
  • Data curation, a term generally defined as a set of activities that includes the preserving, maintaining, archiving, and depositing of data to keep it secure, intact, and accessible for reuse.
  • Many researchers expressed concerns surrounding the ethical reuse of research data. Additional work is needed to establish best practices in this area, particularly for qualitative data sets.
  • Most participants reported feeling adrift when establishing protocols for managing their data and added that they lacked the resources to determine best practices, let alone to implement them. Almost none of the scholars reported that data curation training was part of their graduate curriculum.
  • Perhaps one of the more complicated issues for data curation is the complex life cycle of research data and projects. Data collection may occur throughout the project and change from before it is completed.
  • Scholars may collect data on a phenomenon unrelated to their current project with no clear idea of the potential usefulness of those data. Such data might be integrated with a later project, given away to an interested colleague, or never used at all.
  • It would be helpful to have a way to collect data into a collection space that could be used throughout the project.
  • The researchers held contradictory views about the value of their data. Some wanted to associate their data with publications or to have it available for use in the classroom
  • Few of the researchers thought about long-term preservation of their data, especially those who were early in their career.
  • The academic system offers little or no career reward for preserving one’s data.
  • Data preservation strategies must take into account varied, proprietary, and non-standard data formats, and provide a real-time benefit for the scholar in meeting research goals.
  • Given the lack of infrastructure for sharing and storing data, the social sciences may face similar problems of data loss in documenting social phenomena as researchers begin to work within larger collaborative groups and with larger data sets. Data stored on personal media devices are especially vulnerable to this type of loss, as few scholars have the skills necessary to maintain data over time and across hardware and software platforms. Several of the scholars interviewed reported storing data on legacy systems that may become inaccessible
  • University policies that appropriately address the ethical considerations relating to data sharing and preservation would benefit researchers, administrators, and technologists alike.
  • Researchers hold tremendous amounts of data on personal computers and hard drives, many of which are not backed up adequately. Among the participants, the research data ranged from under 1 GB to multiple terabytes. Data types included various formats of images, video, audio files, data sets, documents, etc.
  • Managing large files presents significant challenges for researchers in that university infrastructures typically do not provide adequate storage space or sufficient bandwidth for data access.  The data may be lost when researchers upgrade their computers or software. Few researchers put more than minimal effort into organizing non-active data or ensuring its continued compatibility with new software or hardware.
  • There is a clear need for libraries to move beyond passively providing technology to embrace the changes in scholarly production that emerging technologies have brought.  
  • The data preservation step must be fully integrated into a scholar’s research workflow. Not only are necessary metadata and other materials much more easily captured while research is in progress, but also there is a real opportunity to streamline research workflows and to provide much needed support. Scholars need help with the technical aspects of managing and preserving data, as well as with basic curation issues (e.g., what to keep and what to delete), and the ethical implications of sharing their data (e.g., what is an appropriate latency period for the data and how does one balance the need to provide meaningful access with the risk of inadvertently exposing confidential participant information).
  • Although some researchers acknowledge that their data could be useful to other researchers, there is little incentive to invest time in archiving or repackaging data sets.
  • Extensive outreach to scholars is necessary to build the relationships that will facilitate data preservation. This is likely to be a slow process initially. Researchers are unlikely to engage with those they do not view as peers.
  • Researchers need additional tools to manage preserved data on their own, and they would benefit from access to professionals who can offer advice on management strategies.
  • Researchers typically align themselves with their disciplines rather than with their institutions; therefore, support models that extend beyond the university are likely to be especially beneficial.
  • Reaching the level of collaboration among universities and the technical interoperability required to capture and preserve a career’s worth of data in the current environment is a challenge.
  • Current data management systems must be fundamentally improved so that they can meet the capacity demand for secure storage and transmission of research data. Integrating the data preservation system with the active research cycle is essential to encourage researcher investment.
  • Researchers are not well positioned to meet the technical and policy challenges without the coordinated support of libraries, information technology units, and professionals who possess both technical and research expertise.
  •  One example concerning the PETRA e+e collider project in Hamburg, Germany; In the more than 25 years since, theoretical insights and computing advancements have made the data valuable once again. However, much of the data have been irrevocably lost to corrupt storage media, lost computer code, and deactivated personal accounts. These early particle physics experiments are unique, as modern colliders operate at higher energy levels and cannot replicate the particle interactions.

Thursday, December 8, 2011

Why don't we already have an Integrated Framework for the Publication and Preservation of all Data Products?

Why don't we already have an Integrated Framework for the Publication and Preservation of all Data Products?   Alberto Accomazzi,et al. Astronomical Data Analysis Software and Systems.  
7 Dec 2011.
Astronomy has long had a working network of archives supporting the curation of publications and data. There are examples of websites giving access to data sets, but they are sometimes short lived.  "We can only realistically take implicit promises of long-term data archival as what they are: well-intentioned plans which are contingent on a number of factors, some of which are out of our control." We should take steps to ensure that our system of archiving, sharing and linking resources is as resilient as it can be.  Some ideas are: 
  1. future-proof the naming system: assign persistent data IDs to items we want to preserve 
  2. provide the ability to cite complete datasets, just as we can cite websites
  3. include a data reference section in academic papers
Curated datasets need to be preserved indefinitely for scholarly purposes.

Tuesday, September 27, 2011

ADS and the Data Seal of Approval – case study for the DCC.

The ADS and the Data Seal of Approval – case study for the DCC.  Jenny Mitcham and Catherine Hardman. Digital Curation Centre website. 2010.  
This page describes the experience of Archaeology Data Service in applying for the Data Seal of Approval (DSA). It provides some practical information about the DSA application process and outlines issues the ADS faced in undertaking the process, and several potential benefits they see from the self-certification.

“When undertaking to curate data for the foreseeable future (and beyond) the concept of ‘trust’ is of paramount importance. Yet in a young discipline such as digital archiving, it is very difficult to demonstrate the potential for longevity of curation.”

The Assessment Manual can be downloaded from the DSA website, which includes details of the 16 guidelines, the minimum requirements, and some guidance notes.  In the spirit of the openness the DSA recommends that the main policy and procedure documents should be accessible the world at large.  One of the benefits mention is it shows to users and depositors that the archive has a set of standards is meeting them.

Thursday, August 25, 2011

Digital Preservation, Digital Curation, Digital Stewardship: What’s in (Some) Names?

Digital Preservation, Digital Curation, Digital Stewardship: What’s in (Some) Names?  Butch Lazorchak. The Signal. August 23, 2011.
We often use “digital preservation,” “digital curation” and  “digital stewardship” interchangeably without thinking about the differences. 
Preservation is defined as keeping something in its origial state.
Curation looks at selection, maintenance, collection and archiving of digital assets in addition to their preservation.

Curation is useful for looking at the entire life of the materials and concentrates on "building and managing collections of digital assets and so does not fully describe a more broad approach to digital materials management.

Stewardship looks at holding resources in trust for future generations which can include both preservation and curation.

Wednesday, August 17, 2011

When Data Disappears.

When Data Disappears. Kari Kraus. The New York Times. August 6, 2011.
A writer said he didn't include digital media in his archive because he felt digital preservation is doomed to fail. “There are forms of media which are just inherently unstable.” It is more difficult, but it is not pointless.  "If we’re going to save even a fraction of the trillions of bits of data churned out every year, we can’t think of digital preservation in the same way we do paper preservation. We have to stop thinking about how to save data only after it’s no longer needed, as when an author donates her papers to an archive. Instead, we must look for ways to continuously maintain and improve it. In other words, we must stop preserving digital material and start curating it."

There are major challenges with digital preservation, but part of it is the amount of data being created.  The world created over 1.8 zettabytes of digital information a year. There will never be enough capacity to save everything if we continue to replicate the practices used to maintain paper archives. In the paper archives model, preservation begins at the end of the life cycle. Data preservation must happen earlier, ideally when the item is created. The decisions about what to save and how to save it must be made early in the life cycle; the data should then be curated, not preserved.  Not all data is worth preserving, either in paper or electronically.  Video games offer an interesting model that may be useful with other types of information.  That model "allows us to see preservation as active and continuing: managing change to data rather than trying to prevent it, while viewing data as a living resource for the future rather than a relic of the past"

Wednesday, August 10, 2011

Hawaiian Heritage sites set to launch on CyArk website.

Hawaiian Heritage sites set to launch on CyArk website. Press release. Hawaii 24/7. August 10, 2011.   
CyArk, with the help of its partners, conducted the field work for the Digital Preservation of three culturally significant Hawaiian sites, which include site animations, photography, panoramas, perspectives, and drawings. This information will showcase oft-overlooked heritage sites and highlight the need for cultural resource preservation.

CyArk is a non profit organization with the mission of digitally preserving cultural heritage sites through collecting, archiving and providing open access to data created by laser scanning, digital modeling, and other state-of-the-art technologies.

Friday, April 2, 2010

Digital Preservation Matters - April 2, 2010

Avoiding a Digital Dark Age. Kurt D. Bollacker. American Scientist. March-April 2010.

Data longevity depends on both the storage medium and the ability to decipher the information

The general problem of data preservation is twofold. The first matter is preservation of the data itself: The physical media on which data are written must be preserved, and this media must continue to accurately hold the data that are entrusted to it. This problem is the same for analog and digital media, but unless we are careful, digital media can be more fragile.

The second part of the equation is the comprehensibility of the data. Even if the storage medium survives perfectly, it will be of no use unless we can read and understand the data on it. Unlike in the analog world, digital data representations do not inherently degrade gracefully, because digital encoding methods represent data as a string of binary digits (“bits”). Because any single piece of digital media tends to have a relatively short lifetime, we will have to make copies far more often than has been historically required of analog media. Like species in nature, a copy of data that is more easily “reproduced” before it dies makes the data more likely to survive.

In order to survive, digital data must be understandable by both the machine reading them and the software interpreting them. There are at least two effective approaches: choosing data representation technologies wisely and creating mechanisms to reach backward in time from the future.

---

A Survey of the Scholarly Journals Using Open Journal Systems. Brian D. Edgar, John Willinsky. Educause Resources. March 4, 2010. [40 p. PDF]

Open Journal Systems (OJS) is an open source, online journal management and publishing platform. This study looks at scholarly communications using the open source software systems. survey to which 998 editors or staff members responded. The results point to how these journals – largely independent, scholar-published titles with roughly half

originating in the developing world – are not otherwise represented. Of the survey, 40 percent published research in the sciences, technology and medicine, 30 percent were social science journals, and 11 percent were in the humanities. 19 percent of the journals in the study were interdisciplinary.

The number of journals using OJS has been growing at an average rate of 81% per year. And the number of new journals that are starting, are using OJS at a rate of 47%. About half the journals using OJS are born digital. OJS looks at the effect that open source tools can have on journal publishing, and adds to the case for rethinking scholarly communication.

---

Ensuring Perpetual Access: establishing a federated strategy on perpetual access and hosting of electronic resources for Germany. The Alliance of German Science Organisations. Final Report in English. March 30, 2010. [177p. PDF.]

Increasing digital content is a challenge for scientific institutions. This study is a basis for a national hosting strategy to “establish and finance sustainable structures for perpetual access as well as long-term preservation for electronic resources.” Research is critical to the economy. Large investments into the research need to be safeguarded and maintained. Any loss can impair research, and ensuring future access is an important challenge. One of the largest gaps is the “provision for perpetual access for e-journals.” Library access via hosting on publishers’ servers is not “sufficiently robust as a single perpetual access solution long-term,” though it may be the immediate approach. Independent perpetual access with partners is needed, such as Portico. There needs to be a “strategy to create an infrastructure for the storage and long-term preservation of digital documents, and which can guarantee perpetual access to licensed commercial publications and retro-digitised library materials.” PDF and XML with the NLM-DTD are becoming a metadata standard for published material.

---

Jhove2-0.6.0 Download. Website. March 19, 2010.

A new alpha release of JHOVE2 is now available for download and evaluation. Some features include:

  • Format identification, validation, feature extraction, and message digest.
  • Recursive processing of directories, file sets, etc.
  • Integration with DROID for file identification.
  • Results formatted as text and XML

---


Friday, March 26, 2010

Digital Preservation Matters - March 26, 2010

Archiving Britain's web: The legal nightmare explored. Katie Scott. Wired. 05 March 2010.

Websites are increasing recognized as being culturally valuable. But there are concerns about the ability to preserve them because of current copyright requirements. The British Library over the past 6 years has archived over 6,000 culturally significant websites. Currently they must contact every copyright holders of these sites, and only have a 24% response rate. Some feel there is a "'digital black hole' in the nation's memory" because of the difficulty in archiving the web sites. There is a proposal to change the law to allow the copy deposit act to include websites. Some look at an opt out option. The BBC has a "no take-down" rule.

---

Canterbury Tales manuscript to be digitized. Medieval news. March 22, 2010.

The University of Manchester Library is planning to digitize the Canterbury Tales manuscript. This is part of a JISC funded project. The Centre of Digital Excellence supports universities, colleges, libraries and museums which lack the resources to digitize important works. In addition to the digitizing work, “they will also be exploring business models for the long term viability of digitisation.”

---

ISO Releases Archival Standards. eContent. Mar 23, 2010.

Two documents from the International Organization for Standardization (ISO) aim to provide guidelines for archiving patient information. "Health informatics-Security requirements for archiving of electronic health records-Principles" and "Health informatics-Security requirements for archiving of electronic health records-Guidelines" look at topics of records maintenance, retention, disclosure, and eventual destruction. Electronic medical data must be stored for the life of the patient; there are legal, ethical, and privacy concerns.

---

Elsevier and PANGAEA Data Archive Linking Agreement. Neil Beagrie. Blog. 03 Mar 2010.

Elsevier and the data library PANGAEA (Publishing Network for Geoscientific & Environmental Data) have agreed to reciprocal linking of their content in earth system research. Research data sets deposited at PANGAEA are now automatically linked to the corresponding articles in Elsevier journals on ScienceDirect. Science is better supported through the cooperation and the flow of data into trusted archives. “This is the beginning of a new way of managing, preserving and sharing data from earth system research.”

---

Duplicating Federal Videos for an Online Archive. Brian Stelter. The New York Times. March 14, 2010.

The International Amateur Scanning League plans to upload the National Archives’ collection of 3,000 DVDs in an “experiment in crowd-sourced digitization” using a DVD duplicator and a YouTube account. This is a small demonstration that volunteers can sometimes achieve what bureaucracies can’t or won’t. the DVDs are all technically available to the public, they are hard to see unless a person visits the archive or pays for a copy. The volunteers duplicate the DVDs then upload them to YouTube, the Internet Archive Web site and an independent server.

---

Uncompressed Audio File Formats. JISC Digital Media. 10 February 2010.

This looks at the main features of uncompressed audio file types, including WAV, AIFF and Broadcast WAV (BWF). “Uncompressed audio files are the most accurate digital representation of a soundwave” but they also take the most resources. Digital audio recording measures the level of a sound wave at regular intervals and records that value as a number. “This bitstream is the ‘raw’ audio data, expressing the sound wave in its closest digital analogue. “ These uncompressed audio file types are ‘wrapper’ formats that take the original data and combine it with additional data to make it compatible with other systems.

The most common is the Waveform Audio File Format (WAV), which is limited to a 4 Gb file size. The European Broadcasting Union created the Broadcast Wave Format (BWF) which is functionally identical to the WAV file except it has an extra header file for metadata. This is a recommended archive format and also has a 4 Gb file size. The European Broadcasting Union has recently added the Multichannel Broadcast Wave Format (MBWF)which combines the RF64 audio format (surround sound, MP3, AAC, etc) with a 64 bit address header and has a file size limit of 18 billion Gb. It is backwardly compatible with WAV and BWF. The Audio Interchange File Format (AIFF) is the native format for audio on Mac OSX.

“The International Association of Sound and Audiovisual Archives (IASA) recommend Broadcast WAV as a suitable archival format, for reasons of its wide compatibility and support, and its embedded metadata capability. For surround-sound or multichannel audio the MBWF format should be used. For archive PCM audio, bit depth should be a minimum of 24-bit, and sample rate a minimum of 48kHz to comply with IASA standards.” If compression is needed, lossless compression, which requires an additional encoding/decoding stage – codec) is the least destructive alternative. Some open-source lossless compression codecs are available, such as FLACC.

---

Court Orders Producing Party to "Unlock" PDF Since Not in a "Reasonably Usable" Form. Michael Arkfeld . Electronic Discovery and Evidence - blog. February 15, 2010.

In this contractual action, the defendants disclosed 11,757-page summary in a PDF "locked" format precluding the plaintiff from being able to edit and or manage the summary without retyping it. The Court found that the defendants' locked format made it "completely impractical for use" and ordered that the defendants "unlock" the files.

---


Tuesday, March 16, 2010

Digital Preservation Matters - March 16 2010

Fending Off Digital Decay, Bit by Bit. Patricia Cohen. The New York Times. March 15, 2010.

This looks at the archival material, including digital, from an author that is on display at Emory University. It highlights what research libraries and archives are discovering, that “born-digital” materials are much more complicated and costly to preserve than anticipated. The “archivists are finding themselves trying to fend off digital extinction at the same time that they are puzzling through questions about what to save, how to save it and how to make that material accessible.” Computers have now been used for over two decades, but their digital materials are just now find their way into archives. The curator said “We don’t really have any methodology as of yet to process born-digital material. We just store the disks in our climate-controlled stacks, and we’re hoping for some kind of universal Harvard guidelines.” The challenges including cataloging the material, acquiring the equipment and expertise to access the data stored on obsolete media. Do they try to save the look and feel of the material or just save the content? The computer editing meant that there are no manuscripts with pages with “lots of crossings-out and scribbling”. The display is providing the “emulation to a born-digital archive” similar to reproducing the author’s work environment. Emory is providing $500,00 to produce a computer forensics lab to do this kind of work. Others are impressed with the emulation, but their focus is storage and preservation of digital content. One center is trying to raise money to hire a to hire a digital collections coordinator. Until then, the digital materials are unavailable to researchers.

---

More on using DROID for Appraisal. Chris Prom. Practical E-Records. March 10, 2010.

The information that DROID supplies is useful but the output not optimally organized for reuse. But by regularizing the DROID CSV output the information became sortable and more useful. DROID was also useful in identifying files that did not use the standard file extension for an application, also to find files that needed attention or need to be converted. And it was very useful in the appraisal process. With it, the major migration problems could be identified and it helped to weed out inappropriate, duplicate, or private content.

---

Data, data everywhere. Economist. February 25, 2010.

The world contains an unimaginably vast amount of digital information which is increasing rapidly. This makes it possible to do many things that previously could not be done but it is also creating a host of new problems. The proliferation of data is making them increasingly inaccessible. The way that information is managed touches all areas of life. The data-centered economy is still new and the implications are not yet understood.

---

Archon™: The Simple Archival Information System. Website. 15 February 2010.

Version 3 of this software has been released. The software is for archivists and manuscript curators. It publishes archival descriptive information and digital archival objects to a user-friendly website. Functionality includes:

· Create standards-compliant collection descriptions and full finding aids using web forms.

· Describe the series, subseries, files, items, etc. within each collection.

· Upload digital objects/electronic records or link archival descriptions to external URLs.

· Batch import data

· Export MARC and EAD records

---

Deluge of scientific data needs to be curated for long-term use. Carole L. Palmer. PhysOrg.com. February 24, 2010.

Data curation is the active and ongoing management of data through their lifecycle. It is an important part of research. Data is a valuable asset to institutions and to the scientific enterprise. Saving the publications that report the results of research isn't enough; researchers also need access to data. Data curation begins long before the data are generated, it needs to start at the proposal stage. Without the data there is the issue of replicating and validating a research project's conclusions. "Digital content, including digital data, is much more vulnerable than the print or analog formats we had before." selecting, appraising and organizing data to make them accessible and interpretable takes a lot of work and expense. "The bottom line is that many very talented scientists are spending a lot of time and effort managing data. Our aim is to get scientists back to doing science, where their expertise can make a real difference to society."

---

Is copyright getting in the way of us preserving our history? Victor Keegan. The Guardian. 25 February 2010.

In theory, future historians will have a lot of information about our age. In reality, much of it may be lost. Much of the information is on web pages, and they have a short life expectancy. The British Library has launched the UK Web Archive, which will guarantee longevity to thousands of hand-picked UK websites. But this is only a small part. “The issue of copyright is a global nightmare for anyone interested in digital preservation.”

---

"Zubulake Revisited: Six Years Later": Judge Shira Scheindlin Issues her Latest e-Discovery Opinion. Electronic Discovery Law. January 27, 2010.

This review of a case that addresses the issues of parties’ preservation obligations. Check here for the full opinion. The case revisits an earlier decision concerning e-discovery, or finding electronic documents, emails, etc, in court cases; obligations; and negligence for failure to keep records correctly. Some statements from the court opinion:

  • By now, it should be abundantly clear that the duty to preserve means what it says and that a failure to preserve records, paper or electronic, and to search in the right places for those records, will inevitably result in the spoliation of evidence.
  • While litigants are not required to execute document productions with absolute precision, at a minimum they must act diligently and search thoroughly at the time they reasonably anticipate litigation.
  • The following failures support a finding of gross negligence, when the duty to preserve has attached: to issue a written litigation hold; to identify all of the key players and to ensure that their electronic and paper records are preserved; to cease the deletion of email or to preserve the records of former employees that are in a party's possession, custody, or control; and to preserve backup tapes when they are the sole source of relevant information or when they relate to key players, if the relevant information maintained by those players is not obtainable from readily accessible sources.
  • The case law makes crystal clear that the breach of the duty to preserve, and the resulting spoliation of evidence, may result in the imposition of sanctions by a court because the court has the obligation to ensure that the judicial process is not abused.


Tuesday, March 2, 2010

Digital Preservation Matters - March 2, 2010

A Guide to Distributed Digital Preservation. Katherine Skinner, Matt Schultz. Educopia Institute. February 2010. [156 p. PDF]

Excellent guide created by MetaArchive, who developed the first private LOCKSS network in 2004. This work examines distributed digital preservation, successful strategies and new models . It will help others to join or establish a private LOCKSS network. It discusses the network architecture, technical and organization considerations, content selection and ingest, administration and copyright practices in the network. A distributed digital preservation system must preserve, not just back-up. The preservation process of contributing, preserving, and retrieving content depends upon the institution’s diligence. Ingested content is preserved not just through replication, but by the caches through a set of polling, voting, and repairing processes. Distributed digital preservation, by definition, requires communication and collaboration across multiple locations and between numerous staff.

The software provides bit-level preservation for digital objects of any file type or format, but it can also provide a set of services to make the preserved files usable in the future, such as normalizing and migrating. The MetaArchive network is a dark archive with no public interface; communication between caches is secure. Organizations collaborating on preserving digital content must examine the roles and responsibilities of members, address essential management, policy, and staffing questions, develop standards, and define the network’s sphere of activity. Ingest, monitoring, and recovery of content are critical steps for preserving the content.

Some interesting quotes from the guide:

  • Paradoxically, there is simultaneously far greater potential risk and far greater potential security for digital collections
  • many cultural memory organizations are today seeking third parties to take on the responsibility for acquiring and managing their digital collections. The same institutions would never consider outsourcing management and custodianship of their print and artifact collections;
  • A great deal of content is in fact routinely lost by cultural memory organizations as they struggle with the enormous spectrum of issues required to preserve digital collections,
  • A true digital preservation program will require multi-institutional collaboration and at least some ongoing investment to realistically address the issues involved in preserving information over time.
  • One of the greatest risks we run in not preserving our own digital assets for ourselves is that we simultaneously cease to preserve our own viability as institutions.

---

Encouraging Open Access. Steve Kolowich. Inside Higher Ed. March 2, 2010.

Conversations about open access to journal articles currently revolve around policy, not technology; about if the content should be made available, not how. “Without content, an IR is just a set of empty shelves.” A new model of repository focuses on giving researchers an online “workspace” within the repository where they can upload and preserve different versions of an article they are working on. The idea is to make publishing articles to the open repository a natural extension of the creative process. This is based on a survey where professors wanted:

  • to be able to work with co-authors easily,
  • to keep track of different versions of the same document, and
  • to make their work more visible
  • all while doing as little extra work as possible.

---

In the digital age, librarians are pioneers. Judy Bolton-Fasman. The Boston Globe. February 10, 2010.

Book review of This Book Is Overdue: How Librarians and Cybrarians Can Save Us All By Marilyn Johnson.

  • Among information professionals, Johnson notes there are librarians and archivists: “Librarians were finders [of information]. Archivists were keepers.’’ But the information revolution is affecting both.
  • The digital age is making possible the creation of searchable databases of archives, but it’s also making information, especially on the Internet, more ephemeral and harder to collect.
  • Information archivists “capturing history before it disappears because of a broken link or outdated software.”
  • in a world where technology moves life at a breathtaking pace, “where information itself is a free-for-all, with traditional news sources going bankrupt and publishers in trouble, we need librarians more than ever’’ to help point the way to the best, most reliable sources.

---

Installing OAIS Software: Archivematica. Chris Prom. Practical E-Records. February 1, 2010.

One of several reports on open source tools the blog author is evaluating to help with ingest, storage, and access processes in archives. This post looks at Archivematica, and he likes the supportable model for facilitating archival work with electronic records. It is a Ubuntu-based virtual appliance which can exist alongside preservation tools on other systems. It can be installed locally and in a variety of ways. Worth looking in to.

---

IBM announces massive NAS array for the cloud. Lucas Mearian. Computerworld. February 11, 2010.

IBM has announced SONAS, an enterprise-class network-attached storage array capable of scaling from 27TB to 14 petabytes under a single name space. It is designed to provide access to data anywhere any time. The policy-driven automation storage software allows an institution to predefine where data is placed, when it is created, where and when it moves to in the storage hierarchy, where it's copied for disaster recovery, and when it will be eventually deleted.