Data growth, data access, and data protection have become intermingled to a degree that makes it difficult to understand management processes that historically were pretty straightforward. Specifically, backup and disaster recovery (DR) have been widely understood and implemented. Due to the enormous growth of data, archive has entered the larger IT discussion, playing an important role in organizing, accessing, and protecting data as it compounds over time, alongside standard backup and DR processes.
By analyzing how large quantities of data may be used and accessed, the roles of each process are much easier to understand.
Historically, backup and disaster recovery have been sufficient for most organizations. To recover from a brief interruption caused by accidentally deleted files, server outages, or power outages, properly implemented data backups let you easily restore data. If a catastrophe strikes, involving a longer-term or more severe interruption, then you implement and test the effectiveness of your disaster recovery plan. These strategies make intuitive sense; they help ensure you can get your data back. The role of archival storage is less obvious.
Backup and Disaster Recovery
Uncertainty is one of the few certainties of life and a governing principle in data center practices. Protecting against the unexpected is one of the significant roles of any IT group â and key to keeping your job and organization intact.
As a result of the evolution of computing, available methods have been adapted to address new needs. Tape initially served as a primary storage medium â storing data that mainframes then accessed, manipulated, and then wrote to tape. As computer memory and disk became more affordable, data and computations were completed online, with tape storing data that was not kept in memory. This spawned the back-up practices of making copies of data in case of an interruption and potential data loss. As computer systems evolved and were more widely used, larger threats were taken into account through disaster recovery planning.
We are now at a point where we have so much data, with the majority of it unstructured, that an additional data management strategy is required: archive.
Beyond Standard Operation and Crisis Management
Data is accessed for reasons beyond daily operations and crisis management. These reasons span the gamut, from accessing huge data sets that are made up of original video and imaging data, intellectual property, older data for business analytics, and enormous caches of raw research data, and more specifically to meet e-discovery requirements and to comply with governmental and organizational regulations. And thatâs just the tip of the iceberg.
Archive Steps In
The need to store infrequently accessed but valuable data has been one of the fallouts of tremendous data growth. It was, for awhile, an issue that was largely avoided â one that some administrators initially attempted to resolve by simply keeping everything on disk. This has proven a poor choice as data growth accelerates. Throwing disk at the problem, and even deduplicating data prior to moving it to disk, simply postpones issues relating to managing data and efficiently storing and retrieving it. Economics of disk-only storage are significant enough to derail the disk-only approach for data centers storing any real quantity of data.
This is where archive fits in. Archival platforms provide proper storage and access methods for data that is infrequently used but retained over long periods, either to meet retention requirements or to provide long-term access to data that may fluctuate in value. Archival systems, which include tape as a key component, add an element of flexibility to the data center, virtualizing content so that administrators can separate capacity demands from performance demands â where disk-only expansion inherently links these. The most obvious example of static data that is useful following latent periods includes news data â anniversary journalism, such as the 10-year anniversary of 9/11 or looking back on the life and contributions of Steve Jobs, requires access to data that may have been dormant for some period. Similarly, when markets (for example, energy) experience ups and downs, research data from years earlier gains renewed value.
Archive is also a fit in environments where enormous quantities of data must be accessible. For example, the media and entertainment industry has used very sophisticated tape-based archives for years. The data is managed and tagged with metadata, and data on tape can be retrieved very quickly â in minutes, rather than hours. Environments using tape in this way typically store tapes in a library so that they can be quickly accessed; this is sometimes referred to as a near-line or near-online environment. Organizations managing large quantities of research data, such as data from ongoing space missions, store data on tape, retrieving it as needed in a near-line environment. The methods these organizations already have in place are now gaining a wider audience as data growth continues unchecked.

The diagram above illustrates the relative data usage pattern of data on disk (not overall capacity usage), from a study completed by the University of California at Berkeley. Extrapolating from the results of the study showing active data use, it is clear that the bulk of the data on disk is typically rarely accessed, making it an excellent candidate for archival storage. Archival storage typically includes tape kept in near-online storage within a library, so that it can be accessed in minutes. The cost savings of tape (compared to spinning disk) and ready availability of data through an archive contribute to the growing popularity of archives across all industries. Implementing an archival system may also greatly reduce the size of scheduled backups.

In many cases, an active archive environment is appropriate â a system that combines applications, open systems, disk and tape hardware to provide users online or near-online access to all data, always. Typically this uses a standard interface, such as a file system interface, and lets anyone (not just administrators) access files. This data is considered online and may reside on disk, tape or both. The archive software pulls the data from the fastest platform (that is, for small files, random access from disk is the fastest if the data is on disk). The key is that the archive handles data transport through user-created policies, and the user can access the data without worrying about the medium on which the data is stored â disk, tape, doesnât matter. The archive application presents the file structure independent of storage method. The user can retrieve data quickly â data on tape that is stored in a contemporary tape library typically requires only a few minutes for retrieval â less if the data is on disk.

Where backup processes focus on making copies of data for security, archive applications typically provide data accessibility, categorizing, and managing unstructured data. As data is migrated into a managed archive through a process termed âingestion,â data is stored and preserved, and, significantly, catalogued. The catalogue itself, also referred to as metadata, lets users search the repository so they can quickly find and retrieve data. Metadata is defined by the end user and typically includes information such as file or data creator, usage statistics, modification information, and other details required for regulatory compliance, as well as content tagging. In this function, the metadata server acts as the gatekeeper to the archive, allowing the data to be stored across multiple storage devices and media types.
Archive applications support long-term retention of data. Many archive applications can also meet an organizationâs e-discovery requirements for rapid retrieval of data, and they further can be set up to automatically purge data from the archive after the compulsory retention requirements have been met. Many archive applications can also track block changes to a file, providing versioning without having to create an individual copy of each iteration of data. This provides the option of non-repudiation.
Archives may be designed to store data online, near-online, or offline. However, one constant across all archives remains: the critical importance of appropriate indexing and metadata. Without proper metadata, information in an archive is inaccessible. Using metadata, the archive can find and restore just the requested information, rather than an entire back-up copy. The speed of this retrieval is a function of the userâs implementation, minimizing time spent locating the data. An additional advantage of metadata-based archives is the ability to monitor and preserve the data within through the use of only two to three copies, rather than the traditional repetitive copying that is a signature of successful back-up strategies.
Protecting Archives as Part of Disaster Recovery Planning
One important best practice, regardless of implementation style, is to maintain at least one, and preferably two, offline copies of all archive data. As with any digital storage, an archive is susceptible to user error, malicious attack, disgruntled employees, and the other unavoidable causes of online data loss.
Best practices advise that a metadata server be backed up as part of nightly backup. However, actual data within an archive may be removed from the traditional back-up regimen. This greatly reduces the size of backups and therefore the length of backup processes while ensuring all organizational data is properly protected. Archives typically maintain multiple copies of a file on each of several storage platforms, and typically have built-in DR features that administrators can set up. These include options such as replication and backup to tape that can be moved off-site.
Itâs All About the Data
The processes used to protect an organizationâs data and continuity have added a layer between nightly backup and disaster recovery: the archiving layer, which makes relatively static data readily available to anyone at any time. This added process addresses the gap between backup, which handles more minor interruptions, and disaster recovery, which enables full organizational rebuilding, while adding the flexibility required by modern data centers.
As data growth continues and data retention requirements increase, static data plays a more prominent role â it must be tagged, monitored and readily available when needed. Data may be accessed for many reasons: business development, organizational continuity, meeting e-discovery/compliance requirements, and a slew of other reasons an organizationâs data serves as its currency.
Christopher Marsh, IT market and development manager at Spectra Logic Corporation, has served in both sales and marketing roles for the organization since 2007. His field of expertise includes thorough knowledge in data storage, disaster recovery, automated tape technology, computer security, and encoding and cryptography. In his current role as market development manager for Spectra, Marsh is responsible for research and analysis of the general IT commercial market. He develops, manages, and oversees strategic relationship management with industry partners to maximize mutual corporate benefits.
