Showing posts with label planning. Show all posts
Showing posts with label planning. Show all posts

Friday, July 8, 2011

Document Capture and Scanning Planning - Part 2

Document Examination and Separation


One of the key steps in preparing for document scanning and capture is to identify how you will separate or split documents.  What is separation and how does it work?  Details below:

For those of you that are new to document management and capture, document separation is the notion of how we can determine when a document begins and ends.  With most simple scanning software, this process is easy.  You load a single document in the feeder, click scan, and when it is done, you name it and save it.  With advanced capture, you can load multiple documents into the feeder, scan them all at once, and use a separation method to split them into individual digital documents.    This is a massive time saver.  Imagine loading 20 individual documents into a scanner one at a time, scanning each individually, and then entering information about each.   Below are some key separation methods any advanced capture suite should have:

Fixed Page Count Separation – This allows you to split based on a certain page count.  So if you scan a stack of 100 two page forms, you will have 50 separate documents in your capture interface.

Barcode Separation – probably the most pervasive separation method is a barcode separator.  Place a sheet with a specific barcode pattern between each document, and you are off to the races.  To give you the most flexibility, applications should support the following enhanced barcode separation methods:

  • Separate on any barcode
  • Separate on specific barcode terms and patterns
  • Separate on barcode type
  • Separate on barcode count
  • Separate on a certain number of barcodes on a page
  • Separate when a barcode changes

You want to make sure your barcode engine supports 1D and 2D barcodes without the purchase of any expensive modules or add-ons, and it should also have a simple feature that lets you split 2D barcodes and identify separation terms.

Patch Code Separation – So what the heck is a patch code?  Just an old school horizontal barcode.  Below is an example.  If you work in the medical field, most medical billing forms will have these on them, and some scanners actually support using patch codes to shift scanner settings during the scanning process.  For flexibility, choose an application that supports patch code separation.

Optical Character Recognition (OCR) Separation – OCR is the process of converting a scanned or imported image into searchable text.  OCR separation searches for a key word, term or phrase on the document, and will recognize that page as the first page in a new document.  This is a preferred method, as you don’t have to kill trees to print cover sheets, and it makes document preparation simple (no inserting separator sheets).  For example, if you are scanning contracts, and you want to split when you find an 8 digit contract number in the right hand corner, this comes in very handy.  There are several key requirements in this feature that are absolutely required in your application to make sure you get high separation accuracy:


  • Scan at 200 or 300DPI and use an app that has image processing software to clean up the page.  Also, your image processing engine must allow processing of imported PDFs and TIFFs if you plan to harvest documents.  Some image correction/processing engines only work with scanners.
  • Insure you capture application allows you to use expression matching (Regular expressions) so you have the utmost flexibility in finding separation patterns.
  • Character sets are key.  These provide the ability to tell the OCR engine the type of characters you are looking for (A-Z, 0-9, etc), so if it misidentifies a character, it auto-corrects the information.
  • Finally, top line applications also allow you to separate when OCR terms change.  So you can look for that contract number, and only split when you find a new one.
Intelligent Character Recognition (ICR) Separation- ICR is the process of converting scanned images of hand printing to text.  This method can be utilized to split pages when certain patterns in hand printing are detected.  Note:  all of the features required to insure accuracy for OCR separation should also be considered if you utilize this method as well.

Document Import and Separation – There are several separation methods that can be key to success if you need to import large volumes of documents, or you want to process documents scanned from copiers, network scanners, or fax machines.  Below is several separation methods required for any document capture from imported files:
  • New File Separation – This method of separation will look at a directory, pick up files, and maintain each new file as its own digital document.
  • Folder-based separation – This is a key method if you are importing documents and want to combine them based on the folder.  One example might be a law firm that has a folder structure of case documents on different subjects for the case and wants to combine each folder into a single PDF file.


Blank Page Separation – I only mention this as I would always, always avoid it unless absolutely necessary, especially if you are scanning in duplex.  Most implementations of this method, unless operated under strict preparation by knowledgeable operators becomes an absolute mess. (Just my humble opinion ;)  )

Separation Scripting – Finally, for those rare and special occasions, you always want a product that has a pre-built scripting interface for customizing the whole process if necessary.  Now let me be clear, not a sales rep “Yeah we can do that” (Which usually means $20,000 in professional services), but a product that has simple hooks into the separation function, that allows you a simple “yes or No” based on some parameter or criteria that anyone with basic scripting skills can write.  When would you use something like this?  Usually for very complex jobs where the original documents cannot be modified, but you need to put some logic in place to spit documents.

The last separation topic I want to cover is something called triggered separation.  Let me set the stage on this one, and describe a process which is near and dear to every accounting manager’s heart, invoices.  So you have a stack of invoices, some single page, some multi-page and you are struck with a dilemma.  If I use barcode separators, and I have 100 single page invoices, do I really have to put 100 barcode separators between them all?  Separation triggers allow you to scan single page and multi-page documents all together.  So in this example, you can stack your singles, and then put separators between your stack of variable length separators.  Put a trigger sheet between the two stacks (this tells the capture software to switch from single page separation to barcode-based separation), and scan the whole stack in one fell swoop.  This is a huge time saver in high volume environments, and can allow you to also build redundant separation logic, so you get the highest accuracy in separation with the least amount of document preparation.  Phewwww.  That was geeky.


Do you really need all of this?  Does separation have to be that complex?  The whole goal here is to have as much as you possibly can in the tool kit to insure you can meet all the capture needs within your organization.  I liken it to buying the a base model with no accessories, and then wishing every day you one or another feature.

So now you have examined your documents, and figured out how to efficiently scan and split.

Wednesday, July 6, 2011

Document Scanning and Capture Planning - Part 1 - Sizing and Storage

Been wanting to do this for quite some time, and finally had some time to sit down and put thoughts together.  I find that many of the scanning and capture implementations lack overall direction, structure and standardization.  I wanted to put together a manual from my experiences, and ask the community to add so we can build a reference for everyone to use.  This will be composed of many parts, including all different topics like storage, hardware, designing your index fields, etc.

Sizing and Storage Planning for Document Management and Scanning



One of the key areas of planning for any scanning/capture implementation is sizing and storage.   Many of the customers we work with have no real grasp on the volume of paper they deal with on a day to day basis, and when they make the migration to digitizing their paper, they are often quite surprised at the amount of paper they push through the system.  Obviously, this can cause some serious issues on many different fronts.   So how do you estimate the amount of paper?  There are several key conversion factors used by the document management industry, as outlined below:

Description
Number of Pages
Storage
1 Scanned Page – 8.5 x 11
1
50KB
1 Scanned Page – 11x17
1
100KB
1 File Cabinet – 4 drawers
10,0000
500MB
1 Box
2500
125MB
1 Linear Inch
100
5MB
1 E Size Engineering Drawing (48x36)
16 – 8.5x11
800KB



This table is a basic planning tool, and can be used as a starting point.  One thing to remember is that these are all standard pages.  Not full image magazine pages, but full text pages.  The other thing to keep in mind is that we have listed for boxes and file cabinets, the average number of pages contained within.  In the imaging world, we deal with images, not pages.  What is the difference?  A page may have 2 sides, which are converted digitally into 2 images.  So effectively, if you have a box with double sided pages you are scanning, you will have to double the storage required.
Some other key factors that can contribute to storage and sizing:

DPI Setting – one of the key questions we always receive is What DPI should I set on my scanner?  For most basic scanning and archive applications, you can set your scanner to 200 DPI.  If you are doing OCR or any type of advanced data extraction, you always want a 300 DPI image for maximum accuracy.  Anything beyond that is just a space killer, will slow down your process and really bloat your files.

Black and White, Greyscale and Color – always use black and white scanning to keep file sizes at an absolute minimum.  Greyscale and color scanning should only be used when absolutely necessary, as file sizes are just crazy.  Below is a table of file sizes for the same letter.  The letter was about 50% page coverage.

Scanning Mode/DPI
File Size
Black and White – 200 DPI
26K
Black and White - 300 DPI
38K
Black and White - 400 DPI
51K
Black and White - 600 DPI
80K
Greyscale – 300 DPI
301K
Color- 300 DPI
577K

Image Processing – image cleanup can significantly reduce file sizes, and it is very important to use this feature whenever you can.  Despeckle, deshade, border removal, etc. will eliminate unnecessary noise in scanned images, and reduce your storage requirement by 10-30% depending on the quality of your documents.

Image Format – There is a lot of misinformation on the market about TIFF versus PDF.  I always hear “We want to store as TIFF because PDFs are just too big.”  Just not the case.  An image scanned to as400 PDF is just a TIFF in PDF clothing (Or a PDF wrapper to be more exact).  The PDF overhead is almost negligible.  The de facto standard in imaging today is rapidly becoming the PDF image with hidden text.  This gives you a nice little file with the pristine image, and converted OCR text in the background.  The text layer adds negligible size to the file.

So now, with all this info, you can estimate volume in images, and then come up with required storage on a monthly, yearly or project basis.

Sunday, January 13, 2008

Document Management and Disaster Recovery

Disaster Recovery is always on the forefront of any solid IT strategy. Companies are becoming so dependent on technology that even the simplest power outage can wreak havoc on business operations. I am always surprised at the lack of attention paper files receive when it comes to the Strategic Disaster Recovery/Business Continuity plan. Companies will spend ten of thousands of dollars on the latest backup and recovery software, offsite data storage and redundant secondary sites, but when asked "What will happen when a fire hits the corporate office and all the paper is gone?", I usually get a blank stare. This is mostly due to the separation of duties within any organization. IT Managers see the data as their responsibility, and go to any length to protect this vital resource. Paper files are almost always managed at the departmental level, by managers who are usually not aware, or educated on disaster recovery and Business Continuity planning.

When examining the overall process, Disaster Recovery and Business Continuity Planning are usually split into separate, but co-dependent processes. Below is a listing of each process and what they include:

Business Continuity Planning (BCP)

  • Plan and Scope organization
  • Business Impact Planning and Analysis
  • Plan Development and Implementation

Disaster Recovery Planning (DRP)

  • Planning process
  • Testing of the plan
  • Recovery procedures

So where does Document Imaging and Document Management play into this process? Paper needs to be a primary focus during the Business Impact Analysis and overall planning exercise. How important are the file cabinets? Can business carry on if all is lost? Is the paper just a redundant copy of existing data? How easily can paper records be recreated?

The whole disaster planning process is a long and arduous task, but organizations need to take into account all their assets to insure business continuity and full operational functionality. Implementing a Document Management and Scanning solution will backup necessary paper files, making sure all required information is available after a disaster.

See the ScanGuru planning section for additional articles and information on planning:

Document Management Planning

Document Management and Data Backup

Tuesday, February 20, 2007

Planning for Document Management and Scanning

When companies are in the planning stages for the implementation of scanning and document management systems, there are several technical infrastructure areas of focus that need to be considered. Most organizations will need to make an investment in their infrastructure to ensure program success and optimal performance. Below are some considerations:

•Storage – Storage planning is critical to provide adequate space and meet future growth requirements for the system. A typical 8 ½ x 11 page will require 50K of storage space. A typical 4 drawer file cabinet contains approximately 10,000 pages, and will require 500 MB of storage on a server. With these benchmarks, you can easily estimate the amount of storage space required. When performing these calculations, make sure and examine your document types to see if they have images, logos, pictures, etc. This will add to the baseline, and increase the amount of storage required. Test scanning of documents is always a good idea to see what the actual size of your image files will be using your scanning hardware of choice.

•Backup – An area often overlooked, a sufficient backup system will be necessary. There are many different philosophies on how to backup a system, but one thing is for sure, larger organizations with a large volume of paper will require a dedicated system for backup. Some options for small offices include CD/DVD, USB Hard drives or a network attached storage device at another location. For large organizations, tape drives and even tape changer systems will be required. Ensure that the device has the ability to backup the entire document repository.

•Network – If you are running your network on 10Mbit hubs, it is probably time to upgrade. Remember, all the facets of a document management system will be transferring large files back and forth between servers, client workstations, MFD scanners and the backup system. You want to invest in the fastest possible network infrastructure to ensure high performance.•Server – That old NT 4.0 server your brother gave you is not going to cut it. Processor speed is not that critical, and any recent server technology will serve well in this environment. Ensure that the server has at least 1GB of memory, and invest in a RAID Disk subsystem for fast access to the files, and redundancy.

•Clients – If you are still running Windows 95, it is time to get up to date. Any modern XP workstation will suffice, and if you have capture workstations that are doing intense document conversion processes (OCR), or are attached to high speed scanners, invest in fast processors and as much memory as you can afford.

Investing the time, resources and capital in a Document Management/Scanning system also requires a modern network to work properly. The investment in modernizing your organizations IT Infrastructure will provide a larger payoff in enhanced system performance, and confidence that the system will be able to grow to its full potential.