Tuesday, March 3, 2015
Oops! Article preserved, references gone
Oops! Article preserved, references gone. Digital Preservation Seeds. February16, 2015.
A blog post concerning the article Scholarly Context Not Found: One in Five Articles Suffers from Reference Rot. References in academic publications justify the argument. Missing references are a significant problem with the scholarly record because arguments and conclusions cannot be verified. In addition, missing or incomplete resources and information will devalue national and academic collections. The Significance method can be used to determine the value of collections. There is currently no robust solution, but a robustify script can direct broken links to Memento. The missing references problem emphasizes that without proper context, preserved information is incomplete.
Saturday, May 11, 2013
ZENODO. Research. Shared.
ZENODO is a new open digital repository repository service that enables researchers, scientists, projects and institutions to share and showcase multidisciplinary research results (data and publications) that are not part of existing institutional or subject-based repositories. The repository is created by OpenAIRE and CERN, and supported by the European Commission. It promotes peer-reviewed openly accessible research; all items have a DOI, so they are citable. All formats are allowed. There is a 1GB per file size constraint. Data files are versioned, but records are not. Files may be deposited under closed, open, embargoed or restricted access.
It is named after Zenodotus, the first librarian of the Ancient Library of Alexandria and father of the first recorded use of metadata, a landmark in library history. ZENODO is provided free of charge for educational and informational use.
Monday, April 22, 2013
Stakeholder Benefits from Research Data Management: new document from Research360 project.
The Research360 Project has released the summary stakeholder benefits analysis from the Research Data Management business case for the University of Bath. The 4 page document is available for download in PDF format.
Industry and private sector partnerships alongside public sector and voluntary sector partnerships are key elements of many university research programmes. Frequently partners sharing their practice, results data and laboratory methodologies can lead to vital knowledge transfer activities, improved services and products, creation of spin-out companies and further investment in the Higher Education sector. A summary list of stakeholder benefits that can arise from research data management in these collaborations. Benefits are listed for:
- university community by its key stakeholder groups:
- academic staff and researchers, students, professional services, and the institution
- external partners:
- industry and commerce, public/voluntary sectors, government, and society
- Improve possibility of success in research funding by addressing any concerns around data management.
- Safeguarding your data against potential loss.
- Support in patent issues such as proof of provenance through improved use of version control.
- Enhanced global reputation through recognition of the quality of research outputs and data infrastructure.
- Attract new collaborators and accelerate deepening of existing relationships.
- Graduate employability increased through university partner connections and student data skills.
- Reduction of risk for sensitive data if data transfer is secure.
- Cost efficiencies from shared data services.
Wednesday, March 27, 2013
Supporting the Changing Research Practices of Chemists.
This report, intended for those who support chemists, including librarians, is about the latest research methods, practices, and information services needs of academics chemists. Chemists need services to make their lives easier and their research groups more productive; this includes minimizing paperwork and administrative tasks. They value academic libraries primarily for the access that they provide to electronic journals and other online resources. Researchers are often frustrated by an inability to share large amounts of data with a collaborator. Few chemists visit the physical library, but they use the library digital collections heavily.
In the survey, fewer than 10% reported a research consultation with a librarian, asked for help with a data management, or asked for assistance on an issue related to publishing in the past year; they rarely reach out to the library to discuss issues or request support. The main search sites for chemists are Web of Knowledge/Web of Science, SciFinder, and PubMed. It would be helpful to have tools to help process all of this information, a pre-scan of announcements from journals, and organize their materials. Electronic Lab Notebooks (ELNs) make it easy to share, archive, and search through past lab notes, but are at risk in the lab. Labs generally do not have good data management infrastructure or proper external support for developing it, especially in sharing and preserving files.
It is difficult for academic chemists to coordinate the recording and preservation of data after the completion of a project. When data are saved, they are often held in unstable or at-risk formats or in formats where no one else can access or interpret them. Sometimes a large amount of potentially useful data is not shared or preserved in any durable way. One chemist invited the library to come and speak to the department about preservation and access. Chemists have a general lack of awareness of effective data curation and preservation. Data management and preservation is time-consuming and rarely straightforward; it requires expert advice and constant monitoring.
The findings:
- Chemists need better support in data management, sharing and preservation.
- Many researchers remain anxious about keeping up with the newest literature.
- They need new tools to stay aware of new research and also serendipitous discovery.
- Chemists require greater support in disseminating their research, including articles, data, and other materials.
Saturday, March 23, 2013
Adding Value to Electronic Theses and Dissertations in Institutional Repositories
This paper looks at the differences with institutional repositories that contain electronic theses and dissertations (ETDs, particularly regarding metadata, policy, access restrictions, representativeness, file format, status, quality and related services. The intent is to improve the "quality of content and service provision in an open environment, in order to increase impact, traffic and usage". This paper shows five ways in which institutions can add value to the deposit and dissemination of electronic theses and dissertation:
- Quality of content. A good IR not only defines a set of standards and criteria for the selection and validation of deposits but also communicates and promotes this editorial policy.
- Metadata. The description of the content and context of the ETD files will make a difference.
- Format. The IR should contain full text, offer different file formats, and have deposit formats are searchable, open, and appropriate for long-term preservation and use of the content.
- Repositories should network and interconnect.
- Provide needed services beyond basic searching, viewing, and downloading. Some possibilities are discussion forums, usage statistics and metrics, citations, Print On Demand in book format, copyright protection or Creative Commons licensing, and preservation.
Monday, February 11, 2013
Sustaining Our Digital Future Institutional Strategies for Digital Content
The shift with digital media in scholarly communications is transformative; data sets, dynamic digital resources, websites, digital collections, crowd sourced or born digital content: there are challenges and opportunities, along with questions about who is responsible for maintaining them, and how to maximize the value of the content.
Some findings:
- project have received support from the host institution, but few have plans for ongoing support.
- There are potential partners on campus, but project leaders do not seek them early
enough when critical decisions are being made. - Digital projects across campuses may be hosted by many groups, which poses challenges for discovery. There is often no single place for users to find digital projects and some projects can too easily slip from view.
- Current funding styles do not support ongoing operation
- Campus-wide solutions are beginning to emerge, but even these tend to address just the basic “maintenance” issues of storage, preservation and access.
- Focus is often on creating new content, with little thought about ongoing efforts to enhance the content or update user interfaces.
- Perform an early and honest appraisal to find which projects are likely to require support after completion:
- Digital content requiring just “maintenance”: plan that the content will be deposited and integrated into some other site, database, or repository.
- Digital resources requiring ongoing growth and investment: These require early sustainability planning, including identifying institutional or other partners and careful consideration of the full range of costs and activities needed to keep the resource vibrant.
- Be realistic in assessing the future needs of the resource at its outset and in continuing support.
- Identify campus partners early on.
- Consider how central your project is to the overall mission of the institution.
- Consider if projects could be drawn together to create a deeper network of support, both for “maintenance” projects and those with the potential to really grow.
- Develop ways to help users find decentralized content and to reach out to content users. These could start as an inventory of all of the digital holdings or common catalogs.
- Determine where scale solutions pay off, where experts are best placed to champion a project, or create common storage, usage and preservation systems for an organization.
- Continue to identify and support ongoing development of the “front-end”, including user needs, interface development, and content enhancement. Pay attention to the changing needs of users and determine what enhancements the digital resource will require.
Sustainability and Use:
- Research data platforms: At some institutions major initiatives are underway to develop research data platforms. The goal of the platform is not just preservation and storage but access and reuse. The first step is to have a platform. From there they can test and refine the service for researchers depositing data sets, library curated collections, and university departments.
- "A coherent digital policy from early review, guidelines on costings and deposit standards, to forecasting what ongoing activities will be needed and who will carry them out, would ideally remove much of the risk of “digital time bombs” while obliging both project leaders and university leaders to take a moment to envision the ongoing impact they want these resources to have, and how to best achieve that."
- Unlike universities (who often play the role of reluctant, passive, or simply unaware host, to a great deal of digital content created by their scholars) museums and libraries tend to be the ones initiating this work and are eager to build and maintain these collections.
- Despite the benefits of centralisation, the mere presence of a catalogue and centralised
repository does not ensure greater usage of or engagement with its holdings.
- Many institutions devote considerable attention to the upfront creation of content, but not nearly as much to its ongoing enhancement or reuse, resulting in collections that are certainly present in the main catalogue, but otherwise exist only as capsules of content, frozen in time.
- Once the project is finished, management of the digital resource is not always clear.
Monday, August 13, 2012
The Problem of Data
Excellent report on data storage, use, and curation. A section contains a snapshot of the current digital data curation education landscape. Below are some long notes and excerpts from the PDF article:
Key Findings
- None of the researchers interviewed for this study have received formal training in data management practices, nor do they expresssatisfaction with their level of expertise.
- Few researchers, especially among those who are early in their career, think about long-term preservation of their data.
- The demands of publication output overwhelm long-term considerations of data curation. Metadata and documentation are of interest only if they help a researcher complete his or her work.
- There is a great need for more effective collaboration tools, as well as online spaces that support the volume of data generated and provide appropriate privacy and access controls.
- Few researchers are aware of the data services that the library might be able to provide and seem to regard the library as a dispensary of goods (e.g., books, articles) rather than a place for research/professional support.
- There is unlikely to be a single out-of-the-box solution that can be applied to the problem of data curation. Instead, an approach is needed that emphasizes working with researchers to identify or build appropriate tools.
- Researchers must have access to adequate networked storage.
- Universities should revise access policies to support multi - institutional research projects.
- Programs should begin early in the researcher career path for the greatest long-term benefit.
- Data curation systems should be integrated with the active research phase (i.e., as a backup, etc).
- Privacy and data access control tools should be developed to manage confidential data. Policies must be developed that support researchers in using these technologies.
- Data curation, a term generally defined as a set of activities that includes the preserving, maintaining, archiving, and depositing of data to keep it secure, intact, and accessible for reuse.
- Many researchers expressed concerns surrounding the ethical reuse of research data. Additional work is needed to establish best practices in this area, particularly for qualitative data sets.
- Most participants reported feeling adrift when establishing protocols for managing their data and added that they lacked the resources to determine best practices, let alone to implement them. Almost none of the scholars reported that data curation training was part of their graduate curriculum.
- Perhaps one of the more complicated issues for data curation is the complex life cycle of research data and projects. Data collection may occur throughout the project and change from before it is completed.
- Scholars may collect data on a phenomenon unrelated to their current project with no clear idea of the potential usefulness of those data. Such data might be integrated with a later project, given away to an interested colleague, or never used at all.
- It would be helpful to have a way to collect data into a collection space that could be used throughout the project.
- The researchers held contradictory views about the value of their data. Some wanted to associate their data with publications or to have it available for use in the classroom
- Few of the researchers thought about long-term preservation of their data, especially those who were early in their career.
- The academic system offers little or no career reward for preserving one’s data.
- Data preservation strategies must take into account varied, proprietary, and non-standard data formats, and provide a real-time benefit for the scholar in meeting research goals.
- Given the lack of infrastructure for sharing and storing data, the social sciences may face similar problems of data loss in documenting social phenomena as researchers begin to work within larger collaborative groups and with larger data sets. Data stored on personal media devices are especially vulnerable to this type of loss, as few scholars have the skills necessary to maintain data over time and across hardware and software platforms. Several of the scholars interviewed reported storing data on legacy systems that may become inaccessible
- University policies that appropriately address the ethical considerations relating to data sharing and preservation would benefit researchers, administrators, and technologists alike.
- Researchers hold tremendous amounts of data on personal computers and hard drives, many of which are not backed up adequately. Among the participants, the research data ranged from under 1 GB to multiple terabytes. Data types included various formats of images, video, audio files, data sets, documents, etc.
- Managing large files presents significant challenges for researchers in that university infrastructures typically do not provide adequate storage space or sufficient bandwidth for data access. The data may be lost when researchers upgrade their computers or software. Few researchers put more than minimal effort into organizing non-active data or ensuring its continued compatibility with new software or hardware.
- There is a clear need for libraries to move beyond passively providing technology to embrace the changes in scholarly production that emerging technologies have brought.
- The data preservation step must be fully integrated into a scholar’s research workflow. Not only are necessary metadata and other materials much more easily captured while research is in progress, but also there is a real opportunity to streamline research workflows and to provide much needed support. Scholars need help with the technical aspects of managing and preserving data, as well as with basic curation issues (e.g., what to keep and what to delete), and the ethical implications of sharing their data (e.g., what is an appropriate latency period for the data and how does one balance the need to provide meaningful access with the risk of inadvertently exposing confidential participant information).
- Although some researchers acknowledge that their data could be useful to other researchers, there is little incentive to invest time in archiving or repackaging data sets.
- Extensive outreach to scholars is necessary to build the relationships that will facilitate data preservation. This is likely to be a slow process initially. Researchers are unlikely to engage with those they do not view as peers.
- Researchers need additional tools to manage preserved data on their own, and they would benefit from access to professionals who can offer advice on management strategies.
- Researchers typically align themselves with their disciplines rather than with their institutions; therefore, support models that extend beyond the university are likely to be especially beneficial.
- Reaching the level of collaboration among universities and the technical interoperability required to capture and preserve a career’s worth of data in the current environment is a challenge.
- Current data management systems must be fundamentally improved so that they can meet the capacity demand for secure storage and transmission of research data. Integrating the data preservation system with the active research cycle is essential to encourage researcher investment.
- Researchers are not well positioned to meet the technical and policy challenges without the coordinated support of libraries, information technology units, and professionals who possess both technical and research expertise.
- One example concerning the PETRA e+e collider project in Hamburg, Germany; In the more than 25 years since, theoretical insights and computing advancements have made the data valuable once again. However, much of the data have been irrevocably lost to corrupt storage media, lost computer code, and deactivated personal accounts. These early particle physics experiments are unique, as modern colliders operate at higher energy levels and cannot replicate the particle interactions.
Sunday, April 29, 2012
Web Archives for Researchers: Representations, Expectations and Potential Uses.
Web archiving is one of the missions of the Bibliothèque nationale de France. This study looks at content and selection policy, services and promotion, and the role of communities and cooperation. While the interest of maintaining the "memory" of the web is obvious to the researchers, they are faced with the difficulty of defining, in what is a seemingly limitless space, meaningful collections of documents. Cultural heritage institutions such as national libraries are perceived as trusted third parties capable of creating rationally-constructed and well-documented collections, but such archives raise certain ethical and methodological questions.
To find source material on the web, some researchers look for non-traditional sources, such as blogs and social networks. Researchers recognize the value of web archives, especially because websites disappear or change quickly. The Internet is no longer just a place for publishing things, “but rather the traces left by actions that people could equally perform in the streets or in a shop: talking to people, walking, buying things... It can seem improper to some to archive anything relating to this kind of individual activity. On the other hand, one of the researchers acknowledges that archiving this material would provide a rich source for research in the future, and thus compares archiving it to archaeology.” Some ask, "How do you archive the flow of time?" New models may be needed. And when selecting an archive, the selection criteria should also be archived, as they may change over time.
Thursday, December 8, 2011
Why don't we already have an Integrated Framework for the Publication and Preservation of all Data Products?
7 Dec 2011.
Astronomy has long had a working network of archives supporting the curation of publications and data. There are examples of websites giving access to data sets, but they are sometimes short lived. "We can only realistically take implicit promises of long-term data archival as what they are: well-intentioned plans which are contingent on a number of factors, some of which are out of our control." We should take steps to ensure that our system of archiving, sharing and linking resources is as resilient as it can be. Some ideas are:
- future-proof the naming system: assign persistent data IDs to items we want to preserve
- provide the ability to cite complete datasets, just as we can cite websites
- include a data reference section in academic papers
Saturday, October 22, 2011
Cite Datasets and Link to Publications
The DCC has published a guide to help authors / researchers create links between their academic publications and the underlying datasets. It is important for those reading the publication to be able to locate the dataset. This recognizes that data generated during research are just as valuable to the ongoing academic discourse as papers and monographs, and in many cases the data needs to be shared. "Ultimately, bibliographic links between datasets and papers are a necessary step if the culture of the scientific and research community as a whole is to shift towards data sharing, increasing the rapidity and transparency with which science advances."
This guide has identified a set of requirements for dataset citations and any services set up to support them. Citations must be able to uniquely identify the object cited, identify the whole dataset and subsets as well. The citation must be able to be used by people and software tools alike. There are a number of elements needed, but the "most important of these elements – the ones that should be present in any citation – are the author, the title and date, and the location. These give due credit, allow the reader to judge the relevance of the data, and permit access the data, respectively." A persistent url is needed, and there are several types that can be used.
Friday, October 7, 2011
More (digital) wake-up calls for academic libraries
The topic was the core business of academic libraries: serving researchers and the scientific research process. There are many changes taking place in the sciences: "zetabytes of data; dynamic, complex data objects that require management; communities and data flows becoming much more important than static library collections, etc." The warning to academic libraries was that if libraries do not develop those services the new researcher needs, someone else will, and then there is no future for the research library. We need a "fundamental transformation process that will affect every aspect of the ‘library’ business." The library needs to provide a repository between the scientific process and IT infrastructure that supports and preserves workflows.
Thursday, September 8, 2011
Research Archive Widens Its Public Access—a Bit
JStor, an organization which maintains link to 1,400 journals for subscribing institutions, is providing free public access to articles published prior to 1923 in the United States or before 1870 in other countries, about 6 percent of its content. In a letter to publishers and libraries, JStor refers to plans for "further access to individuals in the future."
Sunday, September 4, 2011
Institutional Repository and ETD Bibliography 2011
This bibliography has over 600 English-language articles, books, and other works about institutional repositories and theses and dissertations (ETDs). Among other things, it includes digital preservation issues, IR library issues, IR metadata strategies, and institutional open access mandates and policies. Most sources have been published from 2000 through June 30, 2011. The bibliography includes links to freely available versions of included works. It is available as a PDF file.
Thursday, August 11, 2011
Building a Sustainable Institutional Repository
Institutional Repositories are an increasingly important resource and service offered by libraries. Increasing the use of the content is a key to building a sustainable IR. Two organizational types:
1. Structured Content Organization
Organizing content according to its role in the University, which provides a more orderly process of content organization and more efficient metadata.
2. Modular Content Publishing
Creating modules as independent publishing units that work together as a complete and comprehensive publishing system. This uses themed publishing and metadata aggregation.
It is becoming more important for libraries to provide users with the contents and services that are found in institutional repositories.
Friday, May 7, 2010
Digital Preservation Matters - May 7, 2010
The Rocky Mountain News is a good example of what can happen when a newspaper folds. The paper went out of business in February 2009. "All those photos were given to the Denver Public Library and are sitting in a basement in storage. The library can't sell them to me, and they don't have the money to digitize them. So they'll stay in the basement. I spoke to a very nice lady at the library. I said, 'Can they be accessed by the public?' She said, 'Not at this time.' 'Will they ever be digitized?' 'We don't have the funds to do it." Instead, he bought the Denver Post, so he considers the Rocky Mountain News pictures redundant.
- 2008: 48 of 579 URLs, 8.3 %
- 2009: 83 of 579 URLs, 14.3 %
- 2010: 160 of 579 URLs, 27.9 %
Friday, April 16, 2010
Digital Preservation Matters - April 16, 2010
Interesting report about libraries. As the recession continues, Americans turn to libraries in ever larger numbers for access to resources for employment, continuing education, and government services. The local library has become a lifeline of resources, training and workshops. Even in the age of Google, academic libraries are being used more than ever. During a typical week in fiscal 2008, academic libraries in the United States had more than 20.3 million visits, answered more than 1.1 million reference questions, and made more than 498,000 presentations to groups attended by more than 8.9 million students and faculty, increases over the previous years. Over 43% of libraries provide access to locally produced digitized collections.
---
A National Conversation on the Economic Sustainability of Digital Information. Blue Ribbon Task Force on Sustainable Digital Preservation and Access. April 1, 2010. [Silverlight video.]
This page has the agenda and video presentations from A National Conversation on the Economic Sustainability of Digital Information, a recent meeting hosted by the Blue Ribbon Task Force on Sustainable Digital Preservation and Access.
BRTF's Featured Agenda and Presentations:
- Research Data, Daniel E. Atkins, Wayne Clough,
- Scholarly Discourse, Derek Law, Brian Schottlaender,
- Economics of Collectively-Created Content, George Oates, Timo Hannay
- Commercially-owned Cultural Content, Chris Lacinak, Jon Landau
- Economics of Digital Information, William G. Bowen, Hal R. Varian, Dan Rubinfeld
- Summary by Clifford Lynch.
---
How Tweet It Is!: Library Acquires Entire Twitter Archive. Matt Raymond. Blog. Library of Congress. April 14, 2010.
The Library of Congress is digitally archiving every public tweet made since Twitter started in 2006. "Expect to see an emphasis on the scholarly and research implications of the acquisition." Amazing to think what we can "learn about ourselves and the world around us from this wealth of data. And I'm certain we'll learn things that none of us now can even possibly conceive." The Library of Congress has been archiving information from the web since 2000. It now has more than 167 terabytes of web-based information, including legal blogs and political websites.
---
Library of Congress: We're archiving every tweet ever made. Nate Anderson. Ars Technica. April 16, 2010.
Comments about the Library of Congress archiving tweets:
- There's been a turn toward historicism in academic circles over the last few decades, a turn that emphasizes not just official histories and novels but the diaries of women who never wrote for publication, or the oral histories of soldiers from the Civil War, or the letters written by a sawmill owner. The idea is to better understand the context of a time and place, to understand the way that all kinds of people thought and lived, and to get away from an older scholarship that privileged the productions of (usually) elite males."
- Digital technologies pose a problem for the Library and other archival institutions, though. By making data so easy to generate and then record, they push archives to think hard about their missions and adapt to new technical challenges."
---
Aligning Investments with the Digital Evolution: Results of 2009 Faculty Survey Released. Roger C. Schonfeld, Ross Housewright. Ithaka. April 07, 2010. [37p. PDF]
An excellent report for academic libraries especially, Faculty Survey 2009: Strategic Insights for Librarians, Publishers, and Societies, that looks at faculty attitudes towards the academic library, information resources, and the scholarly communications system. A few quotes from the report:
- Faculty most often turn to network-level services, including both general purpose search engines and services targeted specifically to academia.
- Of all disciplines, scientists remain the least likely to utilize library-specific starting points;
- Network-level services are increasingly important for discovery, not only of monographs and journals but archival resources and other primary source collections.
- The library must evolve to meet these changing needs.
- 90% of faculty members view the library buyer role as very important, 71% and 59% now view the archive and gateway roles as very important, respectively. Archiving is the 2nd highest role.
- Despite the reported declines in importance of all the library's roles other than as a buyer, the 2009 study saw a slight rise in perceived dependence on the library
- The declining visibility and importance of traditional roles for the library and the librarian may lead to faculty primarily perceiving the library as a budget line, rather than as an active intellectual partner.
- Faculty members most strongly support and appreciate the library's infrastructural roles, in which it acquires and maintains collections of materials on their behalf.
- Faculty members sense of the significance of long-term preservation of electronic journals has steadily increased over time
- Effective and sustainable models for the preservation of electronic journals must be developed
- Scholars, regardless of field, indicate a general preference that digital materials be preserved.
- Less than 30% of faculty members have deposited any scholarly material into a repository; nearly 50% have not deposited but hope to do so in the future
- Faculty attitudes and practices are at the strategic core. Greater engagement with and support of trailblazing faculty disciplines may help develop the roles and services to serve faculty needs into the future. The institutions that serve faculty must also anticipate them, both to ensure that the 21st century information needs of faculty are met and to secure their own relevance for the future.
Friday, April 2, 2010
Digital Preservation Matters - April 2, 2010
Avoiding a Digital Dark Age. Kurt D. Bollacker. American Scientist. March-April 2010.
Data longevity depends on both the storage medium and the ability to decipher the information
The general problem of data preservation is twofold. The first matter is preservation of the data itself: The physical media on which data are written must be preserved, and this media must continue to accurately hold the data that are entrusted to it. This problem is the same for analog and digital media, but unless we are careful, digital media can be more fragile.
The second part of the equation is the comprehensibility of the data. Even if the storage medium survives perfectly, it will be of no use unless we can read and understand the data on it. Unlike in the analog world, digital data representations do not inherently degrade gracefully, because digital encoding methods represent data as a string of binary digits (“bits”). Because any single piece of digital media tends to have a relatively short lifetime, we will have to make copies far more often than has been historically required of analog media. Like species in nature, a copy of data that is more easily “reproduced” before it dies makes the data more likely to survive.
In order to survive, digital data must be understandable by both the machine reading them and the software interpreting them. There are at least two effective approaches: choosing data representation technologies wisely and creating mechanisms to reach backward in time from the future.
---
A Survey of the Scholarly Journals Using Open Journal Systems. Brian D. Edgar, John Willinsky. Educause Resources. March 4, 2010. [40 p. PDF]
Open Journal Systems (OJS) is an open source, online journal management and publishing platform. This study looks at scholarly communications using the open source software systems. survey to which 998 editors or staff members responded. The results point to how these journals – largely independent, scholar-published titles with roughly half
originating in the developing world – are not otherwise represented. Of the survey, 40 percent published research in the sciences, technology and medicine, 30 percent were social science journals, and 11 percent were in the humanities. 19 percent of the journals in the study were interdisciplinary.
The number of journals using OJS has been growing at an average rate of 81% per year. And the number of new journals that are starting, are using OJS at a rate of 47%. About half the journals using OJS are born digital. OJS looks at the effect that open source tools can have on journal publishing, and adds to the case for rethinking scholarly communication.
---
Ensuring Perpetual Access: establishing a federated strategy on perpetual access and hosting of electronic resources for Germany. The Alliance of German Science Organisations. Final Report in English. March 30, 2010. [177p. PDF.]
Increasing digital content is a challenge for scientific institutions. This study is a basis for a national hosting strategy to “establish and finance sustainable structures for perpetual access as well as long-term preservation for electronic resources.” Research is critical to the economy. Large investments into the research need to be safeguarded and maintained. Any loss can impair research, and ensuring future access is an important challenge. One of the largest gaps is the “provision for perpetual access for e-journals.” Library access via hosting on publishers’ servers is not “sufficiently robust as a single perpetual access solution long-term,” though it may be the immediate approach. Independent perpetual access with partners is needed, such as Portico. There needs to be a “strategy to create an infrastructure for the storage and long-term preservation of digital documents, and which can guarantee perpetual access to licensed commercial publications and retro-digitised library materials.” PDF and XML with the NLM-DTD are becoming a metadata standard for published material.
---
Jhove2-0.6.0 Download. Website. March 19, 2010.
A new alpha release of JHOVE2 is now available for download and evaluation. Some features include:
- Format identification, validation, feature extraction, and message digest.
- Recursive processing of directories, file sets, etc.
- Integration with DROID for file identification.
- Results formatted as text and XML
---
Friday, March 26, 2010
Digital Preservation Matters - March 26, 2010
Websites are increasing recognized as being culturally valuable. But there are concerns about the ability to preserve them because of current copyright requirements. The British Library over the past 6 years has archived over 6,000 culturally significant websites. Currently they must contact every copyright holders of these sites, and only have a 24% response rate. Some feel there is a "'digital black hole' in the nation's memory" because of the difficulty in archiving the web sites. There is a proposal to change the law to allow the copy deposit act to include websites. Some look at an opt out option. The BBC has a "no take-down" rule.
---
Canterbury Tales manuscript to be digitized. Medieval news. March 22, 2010.
The University of Manchester Library is planning to digitize the Canterbury Tales manuscript. This is part of a JISC funded project. The Centre of Digital Excellence supports universities, colleges, libraries and museums which lack the resources to digitize important works. In addition to the digitizing work, “they will also be exploring business models for the long term viability of digitisation.”
---
ISO Releases Archival Standards. eContent. Mar 23, 2010.
Two documents from the International Organization for Standardization (ISO) aim to provide guidelines for archiving patient information. "Health informatics-Security requirements for archiving of electronic health records-Principles" and "Health informatics-Security requirements for archiving of electronic health records-Guidelines" look at topics of records maintenance, retention, disclosure, and eventual destruction. Electronic medical data must be stored for the life of the patient; there are legal, ethical, and privacy concerns.
---
Elsevier and PANGAEA Data Archive Linking Agreement. Neil Beagrie. Blog. 03 Mar 2010.
Elsevier and the data library PANGAEA (Publishing Network for Geoscientific & Environmental Data) have agreed to reciprocal linking of their content in earth system research. Research data sets deposited at PANGAEA are now automatically linked to the corresponding articles in Elsevier journals on ScienceDirect. Science is better supported through the cooperation and the flow of data into trusted archives. “This is the beginning of a new way of managing, preserving and sharing data from earth system research.”
---
Duplicating Federal Videos for an Online Archive. Brian Stelter. The New York Times. March 14, 2010.
The International Amateur Scanning League plans to upload the National Archives’ collection of 3,000 DVDs in an “experiment in crowd-sourced digitization” using a DVD duplicator and a YouTube account. This is a small demonstration that volunteers can sometimes achieve what bureaucracies can’t or won’t. the DVDs are all technically available to the public, they are hard to see unless a person visits the archive or pays for a copy. The volunteers duplicate the DVDs then upload them to YouTube, the Internet Archive Web site and an independent server.
---
Uncompressed Audio File Formats. JISC Digital Media. 10 February 2010.
This looks at the main features of uncompressed audio file types, including WAV, AIFF and Broadcast WAV (BWF). “Uncompressed audio files are the most accurate digital representation of a soundwave” but they also take the most resources. Digital audio recording measures the level of a sound wave at regular intervals and records that value as a number. “This bitstream is the ‘raw’ audio data, expressing the sound wave in its closest digital analogue. “ These uncompressed audio file types are ‘wrapper’ formats that take the original data and combine it with additional data to make it compatible with other systems.
The most common is the Waveform Audio File Format (WAV), which is limited to a 4 Gb file size. The European Broadcasting Union created the Broadcast Wave Format (BWF) which is functionally identical to the WAV file except it has an extra header file for metadata. This is a recommended archive format and also has a 4 Gb file size. The European Broadcasting Union has recently added the Multichannel Broadcast Wave Format (MBWF)which combines the RF64 audio format (surround sound, MP3, AAC, etc) with a 64 bit address header and has a file size limit of 18 billion Gb. It is backwardly compatible with WAV and BWF. The Audio Interchange File Format (AIFF) is the native format for audio on Mac OSX.
“The International Association of Sound and Audiovisual Archives (IASA) recommend Broadcast WAV as a suitable archival format, for reasons of its wide compatibility and support, and its embedded metadata capability. For surround-sound or multichannel audio the MBWF format should be used. For archive PCM audio, bit depth should be a minimum of 24-bit, and sample rate a minimum of 48kHz to comply with IASA standards.” If compression is needed, lossless compression, which requires an additional encoding/decoding stage – codec) is the least destructive alternative. Some open-source lossless compression codecs are available, such as FLACC.
---
Court Orders Producing Party to "Unlock" PDF Since Not in a "Reasonably Usable" Form. Michael Arkfeld . Electronic Discovery and Evidence - blog. February 15, 2010.
In this contractual action, the defendants disclosed 11,757-page summary in a PDF "locked" format precluding the plaintiff from being able to edit and or manage the summary without retyping it. The Court found that the defendants' locked format made it "completely impractical for use" and ordered that the defendants "unlock" the files.
---
Tuesday, March 16, 2010
Digital Preservation Matters - March 16 2010
This looks at the archival material, including digital, from an author that is on display at Emory University. It highlights what research libraries and archives are discovering, that “born-digital” materials are much more complicated and costly to preserve than anticipated. The “archivists are finding themselves trying to fend off digital extinction at the same time that they are puzzling through questions about what to save, how to save it and how to make that material accessible.” Computers have now been used for over two decades, but their digital materials are just now find their way into archives. The curator said “We don’t really have any methodology as of yet to process born-digital material. We just store the disks in our climate-controlled stacks, and we’re hoping for some kind of universal Harvard guidelines.” The challenges including cataloging the material, acquiring the equipment and expertise to access the data stored on obsolete media. Do they try to save the look and feel of the material or just save the content? The computer editing meant that there are no manuscripts with pages with “lots of crossings-out and scribbling”. The display is providing the “emulation to a born-digital archive” similar to reproducing the author’s work environment. Emory is providing $500,00 to produce a computer forensics lab to do this kind of work. Others are impressed with the emulation, but their focus is storage and preservation of digital content. One center is trying to raise money to hire a to hire a digital collections coordinator. Until then, the digital materials are unavailable to researchers.
---
More on using DROID for Appraisal. Chris Prom. Practical E-Records. March 10, 2010.
The information that DROID supplies is useful but the output not optimally organized for reuse. But by regularizing the DROID CSV output the information became sortable and more useful. DROID was also useful in identifying files that did not use the standard file extension for an application, also to find files that needed attention or need to be converted. And it was very useful in the appraisal process. With it, the major migration problems could be identified and it helped to weed out inappropriate, duplicate, or private content.
---
Data, data everywhere. Economist. February 25, 2010.
The world contains an unimaginably vast amount of digital information which is increasing rapidly. This makes it possible to do many things that previously could not be done but it is also creating a host of new problems. The proliferation of data is making them increasingly inaccessible. The way that information is managed touches all areas of life. The data-centered economy is still new and the implications are not yet understood.
---
Archon™: The Simple Archival Information System. Website. 15 February 2010.
Version 3 of this software has been released. The software is for archivists and manuscript curators. It publishes archival descriptive information and digital archival objects to a user-friendly website. Functionality includes:
· Create standards-compliant collection descriptions and full finding aids using web forms.
· Describe the series, subseries, files, items, etc. within each collection.
· Upload digital objects/electronic records or link archival descriptions to external URLs.
· Batch import data
· Export MARC and EAD records
---
Deluge of scientific data needs to be curated for long-term use. Carole L. Palmer. PhysOrg.com. February 24, 2010.
Data curation is the active and ongoing management of data through their lifecycle. It is an important part of research. Data is a valuable asset to institutions and to the scientific enterprise. Saving the publications that report the results of research isn't enough; researchers also need access to data. Data curation begins long before the data are generated, it needs to start at the proposal stage. Without the data there is the issue of replicating and validating a research project's conclusions. "Digital content, including digital data, is much more vulnerable than the print or analog formats we had before." selecting, appraising and organizing data to make them accessible and interpretable takes a lot of work and expense. "The bottom line is that many very talented scientists are spending a lot of time and effort managing data. Our aim is to get scientists back to doing science, where their expertise can make a real difference to society."
---
Is copyright getting in the way of us preserving our history? Victor Keegan. The Guardian. 25 February 2010.
In theory, future historians will have a lot of information about our age. In reality, much of it may be lost. Much of the information is on web pages, and they have a short life expectancy. The British Library has launched the UK Web Archive, which will guarantee longevity to thousands of hand-picked UK websites. But this is only a small part. “The issue of copyright is a global nightmare for anyone interested in digital preservation.”
---
"Zubulake Revisited: Six Years Later": Judge Shira Scheindlin Issues her Latest e-Discovery Opinion. Electronic Discovery Law. January 27, 2010.
This review of a case that addresses the issues of parties’ preservation obligations. Check here for the full opinion. The case revisits an earlier decision concerning e-discovery, or finding electronic documents, emails, etc, in court cases; obligations; and negligence for failure to keep records correctly. Some statements from the court opinion:
- By now, it should be abundantly clear that the duty to preserve means what it says and that a failure to preserve records, paper or electronic, and to search in the right places for those records, will inevitably result in the spoliation of evidence.
- While litigants are not required to execute document productions with absolute precision, at a minimum they must act diligently and search thoroughly at the time they reasonably anticipate litigation.
- The following failures support a finding of gross negligence, when the duty to preserve has attached: to issue a written litigation hold; to identify all of the key players and to ensure that their electronic and paper records are preserved; to cease the deletion of email or to preserve the records of former employees that are in a party's possession, custody, or control; and to preserve backup tapes when they are the sole source of relevant information or when they relate to key players, if the relevant information maintained by those players is not obtainable from readily accessible sources.
- The case law makes crystal clear that the breach of the duty to preserve, and the resulting spoliation of evidence, may result in the imposition of sanctions by a court because the court has the obligation to ensure that the judicial process is not abused.
Tuesday, March 2, 2010
Digital Preservation Matters - March 2, 2010
A Guide to Distributed Digital Preservation. Katherine Skinner, Matt Schultz. Educopia Institute. February 2010. [156 p. PDF]
The software provides bit-level preservation for digital objects of any file type or format, but it can also provide a set of services to make the preserved files usable in the future, such as normalizing and migrating. The MetaArchive network is a dark archive with no public interface; communication between caches is secure. Organizations collaborating on preserving digital content must examine the roles and responsibilities of members, address essential management, policy, and staffing questions, develop standards, and define the network’s sphere of activity. Ingest, monitoring, and recovery of content are critical steps for preserving the content.
Some interesting quotes from the guide:
- Paradoxically, there is simultaneously far greater potential risk and far greater potential security for digital collections
- many cultural memory organizations are today seeking third parties to take on the responsibility for acquiring and managing their digital collections. The same institutions would never consider outsourcing management and custodianship of their print and artifact collections;
- A great deal of content is in fact routinely lost by cultural memory organizations as they struggle with the enormous spectrum of issues required to preserve digital collections,
- A true digital preservation program will require multi-institutional collaboration and at least some ongoing investment to realistically address the issues involved in preserving information over time.
- One of the greatest risks we run in not preserving our own digital assets for ourselves is that we simultaneously cease to preserve our own viability as institutions.
---
Encouraging Open Access. Steve Kolowich. Inside Higher Ed. March 2, 2010.
Conversations about open access to journal articles currently revolve around policy, not technology; about if the content should be made available, not how. “Without content, an IR is just a set of empty shelves.” A new model of repository focuses on giving researchers an online “workspace” within the repository where they can upload and preserve different versions of an article they are working on. The idea is to make publishing articles to the open repository a natural extension of the creative process. This is based on a survey where professors wanted:
- to be able to work with co-authors easily,
- to keep track of different versions of the same document, and
- to make their work more visible
- all while doing as little extra work as possible.
---
In the digital age, librarians are pioneers. Judy Bolton-Fasman. The Boston Globe. February 10, 2010.
Book review of This Book Is Overdue: How Librarians and Cybrarians Can Save Us All By Marilyn Johnson.
- Among information professionals, Johnson notes there are librarians and archivists: “Librarians were finders [of information]. Archivists were keepers.’’ But the information revolution is affecting both.
- The digital age is making possible the creation of searchable databases of archives, but it’s also making information, especially on the Internet, more ephemeral and harder to collect.
- Information archivists “capturing history before it disappears because of a broken link or outdated software.”
- in a world where technology moves life at a breathtaking pace, “where information itself is a free-for-all, with traditional news sources going bankrupt and publishers in trouble, we need librarians more than ever’’ to help point the way to the best, most reliable sources.
---
Installing OAIS Software: Archivematica. Chris Prom. Practical E-Records. February 1, 2010.
One of several reports on open source tools the blog author is evaluating to help with ingest, storage, and access processes in archives. This post looks at Archivematica, and he likes the supportable model for facilitating archival work with electronic records. It is a Ubuntu-based virtual appliance which can exist alongside preservation tools on other systems. It can be installed locally and in a variety of ways. Worth looking in to.
---
IBM announces massive NAS array for the cloud. Lucas Mearian. Computerworld. February 11, 2010.
IBM has announced SONAS, an enterprise-class network-attached storage array capable of scaling from 27TB to 14 petabytes under a single name space. It is designed to provide access to data anywhere any time. The policy-driven automation storage software allows an institution to predefine where data is placed, when it is created, where and when it moves to in the storage hierarchy, where it's copied for disaster recovery, and when it will be eventually deleted.