Thursday, February 5, 2015

Paper Cuts: Why Your Allegiance to Paper Files Could Be Bleeding Your Business


The customer is always right—that's the old adage. These days it’s even more than that, though. The customer determines whether or not your business thrives. A bad review or a negative testimonial online can have serious implications for any company, and it can be difficult to get rid of the bad press once it’s been posted. And unfortunately, bad reviews are much louder than good ones—the White House Office of Consumer Affairs reports that news of poor customer service reaches twice as many people as does praise for good customer service. Research also shows that it takes 12 positive experiences to make up for just one unresolved negative experience. 

The best defense against unhappy customers is pretty obvious: provide great customer service. But it’s not just as simple as that. Try as you might, if your business hasn’t yet made the move toward a paperless or reduced-paper office, you might be missing out on opportunities to be there for your customers.
Unfortunately, research from the Gartner Group shows that professionals spend 50% of their time searching for information, with an average of 18 minutes spent to locate a document. Additionally, an estimated 15% of paper documents end up misplaced or misfiled. And those misplaced documents can cost a lot more to find and replace than you might think. See some of the costs associated with paper documents here
It’s hard to imagine that, given how much time employees spend on these menial tasks, they are really serving customers to their fullest potential. 

The good news is that it doesn’t have to be this way! You can stop paper’s takeover of your office and get your employees back to doing what they’re good at: making your customers happy. All you have to do is take out the part of the equation that’s costing your employees their valuable time—paper.

But where do you start? It’s got to be a lot of work to get all of your company’s existing files into digital form, right? It’s actually a surprisingly simple process if you choose the right tools. Let us show you how. Visit psigen.com and take a look around, or contact one of our experts here

Friday, March 21, 2014

The Cloud: 5 Things to Do Before Adopting Cloud ECM

Don't Let the Cloud Kill Your Network



Companies have been slow to adopt cloud-based ECM for a variety of reasons: security, perceived lack of control and lack of integration.  Scanning high image volumes to the cloud can kill your network, and cause major issues.  Take these 5 steps to make sure a smooth roll out:

  1. Assessment is key.  Doing a file assessment and analysis should be done immediately.   Take a deep dive into each of your departments, and figure out their scanning and capture needs.  Does your legal department want to scan 500 page documents?  Is back-scanning of file cabinets going to be a major portion of the project?  Does marketing want to scan full-page color?  Key areas to be identified are: large document scanning, color requirements, and high volume areas.  For more information on planning and assessment see here:   Scanning Planning
  2.  Check your internet bandwidth, and monitor.  IT involvement from a monitoring perspective will be key to ensure you proper bandwidth to support your scanning efforts.  Batch uploads from large file scanning can kill bandwidth quickly, and create a user revolt.  Proof of concept and single department implementations can give great insight into network impact, and provide some great stats for follow on phase roll outs.
  3.  Check your device settings, and control them.  Most scanners and copiers today will scan in full color if you let them.  File sizes vary to the extreme between black and white, grayscale and color.  Along with color settings, DPI should be controlled, and in most cases 200 DPI black and white is sufficient for most organizations needs.  Nothing kills a party like a 500MB color scan!!  Tips for Scanning Copier settings:  Copier Settings that Kill
  4. Check your server side settings.   Does your ECM System  set file upload limitations.  Make sure from your file assessment that you will be able to handle all file sizes required.  If you cannot control these settings, or your provider will not change them, make sure you use a capture technology that can perform file splitting for you .
  5. Timing can be key.  Depending on your requirements, it may be necessary to control large uploads.  For example, some customers have chosen to do their back scanning and large uploads during off hours / weekends so as to not impact daily operations.  Others will coordinate with a 3rd party scanning service to perform all their high volume scanning off site, with a planned, controlled upload during off hours.
Anything I missed?  Comments from the trenches?  Please post your comments.

Wednesday, September 11, 2013

SharePoint Scanning Case Studies

Ran into a great site with some really cool Scan to SharePoint Case Studies.  This company is out of the UK, Datafinity, and has several deployments where customers are scanning and capturing documents into SharePoint libraries.  Here are some summaries:

Kepak Group is a young, professional and dynamic business that has grown into one of Europe's leading food processing companies employing over 2000 people in nine manufacturing facilities across Ireland and the UK. Kepak installed PSI:Capture to manage the scanning and indexing of a growing volume of personnel files that needed to be stored and managed. PSI:Capture links to payroll and HR systems to retrieve index data, converts the HR files into text-searchable PDFs and transfers the files to SharePoint 2010.  “PSI:Capture has enabled us to scan large volumes of documents into pre-defined structures in SharePoint with the use of very simple drop-down menu options and links to our HR system” said Aine Black, HR Manager, Kepak Group. Read full case study

Haulfryn Group Ltd, an operator of holiday and residential mobile home parks across England and Wales, decided to deploy PSI:Capture Enterprise as their SharePoint 2010 Document Capture solution.  Using PSI:Capture in conjunction with Kodak document scanners, they now have an end to end capture solution that provides unmatched speed and automation, along with a simple, yet powerful user interface.  “PSI:Capture has made our whole scanning process robust,” said Stephen Lattimore, Business Process Manager for Haulfryn.  “We can now quickly reference our documents in SharePoint 2010 for audit and service.”  Along with the current scanning process, Haulfryn plans to add the processing of survey forms and other documents in the near future. Read full case study

The Fire Brigades Union, headquartered in Kingston-upon-Thames, needed a way to store large volumes of paper documents in their newly deployed Microsoft SharePoint 2010 document management system. They evaluated several scanning and OCR products available on the market before choosing PSI:Capture, because of its ease of use, quick implementation and unparalleled interface to SharePoint. The Union now scans many thousands of documents a day which are converted into text-searchable PDFs and stored in SharePoint providing instant access to all paper information for the Union staff located throughout the UK. Read full case study

Isos Housing, a housing association headquartered in Newcastle upon Tyne, is responsible for the day-to-day management of almost 12,000 homes across the North East, from Berwick in the north down to Stockton in the south, and across to Cumbria in the west. The company adopted PSI:Capture to enable them to automate the scanning and storage of invoices and other accounting documents in their Microsoft SharePoint system. PSI:Capturereads unique barcode references created from their accounting system, Open Accounts, to index and organise documents in SharePoint for quick and easy access by staff in their four offices across Northumberland.

Friday, August 30, 2013

Featured Webinar: Your Profit is in Danger

Your Profit is in Danger
Join us for a Webinar on September 10
Space is limited.
Reserve your Webinar seat now at:
https://www1.gotomeeting.com/register/778535032
This joint webinar with PSIGEN and OPEX will focus on how to improve document scanning efficiency through a combination of PSIGEN PSI:Capture Enterprise and OPEX hardware.  See how you can reduce prep time and save on labor, improving your margins and driving higher profits.

Title:
Your Profit is in Danger
Date:
Tuesday, September 10, 2013
Time:
10:00 AM - 11:00 AM PDT

After registering you will receive a confirmation email containing information about joining the Webinar.

System Requirements
PC-based attendees
Required: Windows® 8, 7, Vista, XP or 2003 Server
Mac®-based attendees
Required: Mac OS® X 10.6 or newer
Mobile attendees
Required: iPhone®, iPad®, Android™ phone or Android tablet

Wednesday, July 11, 2012

Mobile Capture? Really?

The Document Management industry is all about mobile capture right now. Really? Taking pictures of documents, page by page, with a tablet/smart phone camera. Some of the biggies in the industry are spending huge amounts of money promoting the cause, and building complex infrastructures and image processing to handle these types of images. There are a number of new startups, like StratusFlow, that are focusing on solving the key problem through the cloud.  Want to see a simple solution? Video below uses Microsoft SkyDrive, an iPad and PSI:Capture on the backend to read barcode photos and process the data.

 

Thursday, May 31, 2012

How do you want to find your documents?


Document Capture Drives Search
One of the first stages in planning for any scanned image repository is to ask the question: How do you want to find your documents?  Theories vary on best practices, but here are a few tips when designing a document capture implementation for any ECM system:
  1. Limit your number of fields to 5 or less. So many times i see document scanning customers use way to many fields during capture.  The more fields you have, the more time for end users to index their documents, and the more chances fields will get skipped.  Take the time to interview the end users and truly find how they need to search for their documents.
  2. Always use a date.  Dates are the ultimate filter that can be a life saver when searching for that needle in a haystack in a scanned document repository.  Invoice date, purchase order date, contract date, etc. give you the power to narrow down your search results to a specified period and can be a huge help in audit based searches or searches for legal support.
  3. Use automation to reduce indexing time.  Document capture applications provide automation and efficiency, and can reduce end user keying requirements on documents.  Strong, accurate OCR technology, and Advanced Data Extraction (ADE) are absolutely required.
  4. Ensure your technology has a QA step.  If you are going to go to all the trouble of scanning, capturing and migrating documents to a repository, make sure you can check your work.  Misfiling a document can a painful experience.
  5. Full text search is the insurance policy.  Always, I repeat always, convert your scanned documents to a searchable format, PDF Image with Hidden text.  This will allow for granular searches beyond your index fields/columns, and can help you in the "find a needle in the haystack" tasks.  But do not, I say, do NOT rely on full text search as your primary search method.  Full text does not let you sort by specific document focused dates, cannot let you do range based searches on specific criteria, and restricts sorting and viewing in most repositories.
Just a few tips when designing your document scanning index fields.

Monday, April 30, 2012

Oooops. Did someone backup the paper?

If you  look at the headlines over the past few years, you cannot help but notice the number of natural disasters that have occurred.  In my conferences with IT and Departmental Management, I always pose the question when discussing business continuity or disaster planning: Do you have a plan for your paper?   Just about every company has implemented some type of plan for backing up their important digital files.  Some go to the extreme with data snapshots that can be recovered from multiple locations.  But companies typically don't take the same strategy with their paper assets.  The good ole file cabinet, the protector of all things paper will provide protection, right? Companies need to take a good hard look at their paper, and assess the business impact should disaster destroy their file room.  Backing up your paper nowadays is not hard, nor expensive when compared to the legal implications and time it would take to reproduce (if possible) contracts, customer files, sales records and the like. Any paper backup plan involves a concept i call Bridging the Gap (BTG).  BTG involve hardware and capture software to digitize and build the bridge to the digital world, and then a repository on the "other side" to house the records and make search and retrieval simple.  The repository can be as simple as a set of named network folders, or as complex as a true ECM system like MS SharePoint.  Take the initiative and backup your paper today.

Monday, October 17, 2011

Document Scanning and Capture Planning - Part 4 - Document Scanning Models


Document Scanning Models

After doing some planning on the hardware types and document scanning volumes, the next step would be to examine what type of model you need to deploy.  There are typically 3 standard  models for document scanning and capture: Centralized, De-centralized and Distributed. 
Each model has its own pros/cons, and below I will examine each, and dive into some detail.
Centralized
Ah, the centralized model.  Some call this old school scanning and capture, as for many years, this was the only way to get the job done, and convert your paper to digital form.  This model provides a centralized scanning center to provide mass conversion for the organization.  The operation can be run by in house personnel, be managed by a services provider in house, or be outsourced to a scanning service bureau.  It requires high volume/high speed hardware, and typically utilizes advanced capture software to allow for the utmost in automation and efficiency.  The software and hardware operators are typically highly trained, and there are usually only a few of them.  Paper and/or digital media is shipped to the centralized location and processed through a set, standardized capture workflow.
Centralized Pros
  • Easily standardized process due to a limited number of skilled/trained scan operators
  • High speed hardware/software results in minimal processing time once paper is received
  • Centralized reporting and control of overall process
  • No loading on WAN infrastructure
  • Centralized backup and restore
Centralized Cons
  • Usually a high time delay for availability of documents
  • High cost due to shipping of documents
  • High maintenance costs
  • High training costs to bring on new operators
  • Disaster recovery planning issues if centralized site is down
  • Operators are typically not knowledgeable in the documents they are indexing
Decentralized
Over time, as bandwidth and scanning hardware/software prices went down, the obvious move was to decentralize the whole scanning and capture process.  This move placed scanning in the branches, and allowed the whole document capture process to be performed by those who had working knowledge of the documents.  Smaller, desktop class hardware could be used, and most capture companies made batch scanning and upload to the centralized repository simple to accomplish.
Decentralized Pros
  • Scan operators are well versed in the documents they scan
  • Documents are available almost immediately
  • No shipping or transfer costs for documents
  • Branch control of the whole scanning process
Decentralized Cons
  • Standardization can be an issue
  • No centralized control or reporting
  • WAN Bandwidth consumption can be high
  • Licensing costs can be high depending on software utilized
Distributed
The advance of network-based scanning devices and the lowering of bandwidth pricing led to the newest model, the Distributed Model.  Distributed Scanning allows for just about anyone in the organization to walk up to a network scanning device/scanning copier/fax machine and send documents to a repository.  The devices are typically multi-faceted, and along with repository integration, can provide scan to network folder, FTP and email.  Collaborative back-end systems, like Microsoft SharePoint, lend themselves nicely to this model, as they allow anyone to participate in a Document Workspace.
Distributed Pros
  • Put scanning in the hands of everyone in the organization
  • Provides a great launching pad for collaborative solutions
  • Simple, easy to use interfaces allow for minimal training and quick adoption
  • Capture and indexing is now in the hands of the true document owner
  • One-to-many solution provides a single device to service many users
Distributed Cons
  • Lack of standardization without software addition
  • Security and document control can be major issues
  • Bandwidth from smaller branches can be a problem with larger scans
  • Lack of hardware integrations with back-end systems
So, most organizations today are combining the above models to create a Hybrid Scanning and Capture solution, and leveraging all the strengths together to minimize the weaknesses of any one model.   Another strategy is to tie scanning models to specific business processes, as most lend themselves nicely to specific scanning and capture workflows.

Hardware and Choosing Your Scanning Model


Most organizations will choose their model to leverage their existing hardware investment, but this can be lead to decisions that seem good at the time, but if deeper examination occurs, it can make sense to realign hardware with the best model.  Take for example, a company that instantly leans toward a distributed model, and attempts to leverage their copier fleet that is currently under lease.  If you examine the part of this guide that covers scanning hardware, copiers will not always fit for the type of scanning you need to perform.  Take for example a branch accounting department that is looking to scan receipts or check stubs.  Will the copier perform well with mixed original sizes?  Just a word of caution to examine the paper, workflow, and document types to get the best feel and adapt the best model.

Tuesday, August 2, 2011

Document Scanning and Capture Planning - Part 3 - Scanning Hardware

Now that I have covered Sizing and Storage in Part 1, and Document Separation in Part 2, now we can start to take a look at scanning hardware.  There are several key questions you need to answer:  Can I use pre-existing hardware such as copiers or fax machines?  Do I need a dedicated scanner?  If I choose to buy a scanner, what features/characteristics are important?

Some may argue you need to decide on a scanning model before you dive into hardware (distributed, centralized, or decentralized), but I will cover this in the next section.

So let’s start with a key question:

Scanning Copier or Dedicated Scanner??

Scanning Multifunction Peripherals (MFPs/copiers) have become standard in most offices. I receive the same question all the time from prospects and customers: Can’t I just use my copier for scanning? In many cases, for a typical office, with typical documents, a copier can be an appropriate component to any scanning solution. As offices become more complex in the way they handle their documents, or they expand their scanning efforts to other departments, dedicated scanners are usually required to achieve the desired result.

Below are some interesting statistics provided by InfoTrends:

· 65 % of office workers use digital copiers/MFPs
· Over 50% use the “scan” feature daily
· 71% expect scanning requirements to increase from year to year
· 72% believe it is necessary to view images before processing
· 36% will require dedicated scanners versus MFP devices
· 36% believe they will need both scanners and MFPs

So what are the benefits/drawbacks to scanning with both types of devices? Below is a summary:

Benefits of MFPs as scanners:

  • Leverage your existing investment in the MFP
  • Most copier maintenance plans do not charge for scans, so you get “free” maintenance for the scanning function (no print/copy, no click charge)
  • MFP manufacturers are really focusing on scanning capabilities: fast speeds, better quality and enhanced drivers, etc.
  • Network scanning functions:
  • Scan to email
  • Scan to Windows Folders
  • Scan to FTP
  • One-to-Many relationship: all workers can use one device.

Drawbacks of MFPs:

  • Contention – copying, scanning and printing may cause “a line at the copier”
  • Poor performance with differing paper sizes
  • Lack of color dropout (Scanning blue or black backgrounds will result in a black page)
  • Lack of image correction capabilities (auto deskew, despeckle, black border removal, streak removal, etc.)
  • Small Document Feeder sizes (50 – 100 pages)
  • On average, file sizes are 10-20% larger
  • Duplex scanning/DPI increase greatly slows down rated speed
  • Black and White scanning only on some models

Benefits of Dedicated Scanners:

  • Convenience – scan at your desk
  • Duplexing does not slow down scanner
  • Color dropout
  • Superior image quality due to enhancement features
  • Ease in handling differing paper sizes/types
  • Larger document feeder selections (up to 1000+ pages)
  • Smaller file sizes
  • Ability to preview scanned documents at scan time

Drawbacks of Dedicated Scanners:
  • One to One relationship – directly connected to PC
  • Additional Maintenance costs

Above are all the pluses and minuses, but in a nutshell, when should you use a dedicated scanner?

  • Scanning 50+ documents per day
  • Workers that are constantly scanning throughout the day
  • Mixed paper sizes, weights and colors
  • Poor quality, older documents or when image enhancement is required
  • OCR or ICR applications
  • High volume copying and printing environments
  • Large Document scanning
  • High security environments

Now that you have an idea of the pros/cons of both types of scanning devices, now let’s take a look at the different features of scanning devices, and what to look for when purchasing a dedicated scanner.


Scanning Speed

Scanning speed is a main area of focus when researching scanning hardware. A scanner’s speed is usually directly proportional to its price, but you have to ask yourself one question: How long do you have to accomplish your scanning tasks? If you buy that cheapo scanner at an office products store that scans at 8 pages per minute, good luck in getting those 10 file cabinets scanned. Another note to mention is that all the manufacturers rate their scanner speeds at 200 DPI. If you need high quality images, or are performing OCR, 300 DPI will probably be necessary. This will significantly slow down your scanning speed, as will color scanning and duplex (2-sided) scanning on some models.

Document Feeder Capacity

The document feeder provides you the ability to load anywhere from 1-1000+ sheets into the scanner. The feeder capacity you require all depends on the volume of paperwork you are scanning, and if you are using an intelligent capture application that provides the ability to use separator sheets to split documents automatically. If you are a Law Firm that routinely scans 200 page documents, then that is a good starting point for your feeder size requirements. This allows you to load your documents, and then let the scanner do the work.

Another focus area related to the feeder is the maximum and minimum paper sizes. If you intend to scan legal size paper or insurance cards, make sure the scanner can handle them.

Daily Duty Cycle

The Duty Cycle (DC) is a rating of the scanner’s durability, and defines just how much paper you can feed through the hardware in a day. If you are scanning 3000 pages per day, you do not want to buy a small desktop scanner with a DC of 750. What happens if you exceed this number? Nothing to begin with, but as time goes on the wear and tear on the unit will begin to show in the form of jams, miss feeds, skewing, etc. This number is also tied to the replacement of consumables (rollers and pads). If you continually exceed the DC, you will more than pay for a higher level scanner in consumables over time, and your maintenance costs may go way up.

Scanning Mode

Most scanners nowadays can scan both sides of your document, but there are still some lingering models that will only do simplex scanning. Also, if you have the requirement to scan color documents, ensure that color scanning is supported.

Warranty and Service

All warranties are not created equal. Some scanner manufacturers provide “depot” type service where you have to ship your scanner for warranty service. Others will provide onsite warranty service for a specified period of time. Along with this, the time period on the warranty also varies everywhere from 30 days, to a full year. Scanner service is a separate purchase, and in some cases, can be a shock to the purchaser. A basic service plan on a mid-range scanner can cost over $1000 per year. Get an advanced plan that provides Preventative Maintenance visits, and you could be in the $1500 - $2000 range, depending on your model. Get all the details up front, and some manufacturers will provide multi-year discounts on service.

Image Processing

Definitely investigate the image processing software that comes bundled with your scanner.  This software will improve the quality of your images, remove shading, borders, etc.  Many of the manufacturers now provide third party image processing software (Kofax VRS), but several have their own built into their drivers.  Most capture software also has built in image processing components as well.

So hopefully this will answer the majority of your questions on hardware.  Remember, hardware is just part of the overall capture solution.  Follow on articles will cover information on software selection and required features.

Friday, July 8, 2011

Document Capture and Scanning Planning - Part 2

Document Examination and Separation


One of the key steps in preparing for document scanning and capture is to identify how you will separate or split documents.  What is separation and how does it work?  Details below:

For those of you that are new to document management and capture, document separation is the notion of how we can determine when a document begins and ends.  With most simple scanning software, this process is easy.  You load a single document in the feeder, click scan, and when it is done, you name it and save it.  With advanced capture, you can load multiple documents into the feeder, scan them all at once, and use a separation method to split them into individual digital documents.    This is a massive time saver.  Imagine loading 20 individual documents into a scanner one at a time, scanning each individually, and then entering information about each.   Below are some key separation methods any advanced capture suite should have:

Fixed Page Count Separation – This allows you to split based on a certain page count.  So if you scan a stack of 100 two page forms, you will have 50 separate documents in your capture interface.

Barcode Separation – probably the most pervasive separation method is a barcode separator.  Place a sheet with a specific barcode pattern between each document, and you are off to the races.  To give you the most flexibility, applications should support the following enhanced barcode separation methods:

  • Separate on any barcode
  • Separate on specific barcode terms and patterns
  • Separate on barcode type
  • Separate on barcode count
  • Separate on a certain number of barcodes on a page
  • Separate when a barcode changes

You want to make sure your barcode engine supports 1D and 2D barcodes without the purchase of any expensive modules or add-ons, and it should also have a simple feature that lets you split 2D barcodes and identify separation terms.

Patch Code Separation – So what the heck is a patch code?  Just an old school horizontal barcode.  Below is an example.  If you work in the medical field, most medical billing forms will have these on them, and some scanners actually support using patch codes to shift scanner settings during the scanning process.  For flexibility, choose an application that supports patch code separation.

Optical Character Recognition (OCR) Separation – OCR is the process of converting a scanned or imported image into searchable text.  OCR separation searches for a key word, term or phrase on the document, and will recognize that page as the first page in a new document.  This is a preferred method, as you don’t have to kill trees to print cover sheets, and it makes document preparation simple (no inserting separator sheets).  For example, if you are scanning contracts, and you want to split when you find an 8 digit contract number in the right hand corner, this comes in very handy.  There are several key requirements in this feature that are absolutely required in your application to make sure you get high separation accuracy:


  • Scan at 200 or 300DPI and use an app that has image processing software to clean up the page.  Also, your image processing engine must allow processing of imported PDFs and TIFFs if you plan to harvest documents.  Some image correction/processing engines only work with scanners.
  • Insure you capture application allows you to use expression matching (Regular expressions) so you have the utmost flexibility in finding separation patterns.
  • Character sets are key.  These provide the ability to tell the OCR engine the type of characters you are looking for (A-Z, 0-9, etc), so if it misidentifies a character, it auto-corrects the information.
  • Finally, top line applications also allow you to separate when OCR terms change.  So you can look for that contract number, and only split when you find a new one.
Intelligent Character Recognition (ICR) Separation- ICR is the process of converting scanned images of hand printing to text.  This method can be utilized to split pages when certain patterns in hand printing are detected.  Note:  all of the features required to insure accuracy for OCR separation should also be considered if you utilize this method as well.

Document Import and Separation – There are several separation methods that can be key to success if you need to import large volumes of documents, or you want to process documents scanned from copiers, network scanners, or fax machines.  Below is several separation methods required for any document capture from imported files:
  • New File Separation – This method of separation will look at a directory, pick up files, and maintain each new file as its own digital document.
  • Folder-based separation – This is a key method if you are importing documents and want to combine them based on the folder.  One example might be a law firm that has a folder structure of case documents on different subjects for the case and wants to combine each folder into a single PDF file.


Blank Page Separation – I only mention this as I would always, always avoid it unless absolutely necessary, especially if you are scanning in duplex.  Most implementations of this method, unless operated under strict preparation by knowledgeable operators becomes an absolute mess. (Just my humble opinion ;)  )

Separation Scripting – Finally, for those rare and special occasions, you always want a product that has a pre-built scripting interface for customizing the whole process if necessary.  Now let me be clear, not a sales rep “Yeah we can do that” (Which usually means $20,000 in professional services), but a product that has simple hooks into the separation function, that allows you a simple “yes or No” based on some parameter or criteria that anyone with basic scripting skills can write.  When would you use something like this?  Usually for very complex jobs where the original documents cannot be modified, but you need to put some logic in place to spit documents.

The last separation topic I want to cover is something called triggered separation.  Let me set the stage on this one, and describe a process which is near and dear to every accounting manager’s heart, invoices.  So you have a stack of invoices, some single page, some multi-page and you are struck with a dilemma.  If I use barcode separators, and I have 100 single page invoices, do I really have to put 100 barcode separators between them all?  Separation triggers allow you to scan single page and multi-page documents all together.  So in this example, you can stack your singles, and then put separators between your stack of variable length separators.  Put a trigger sheet between the two stacks (this tells the capture software to switch from single page separation to barcode-based separation), and scan the whole stack in one fell swoop.  This is a huge time saver in high volume environments, and can allow you to also build redundant separation logic, so you get the highest accuracy in separation with the least amount of document preparation.  Phewwww.  That was geeky.


Do you really need all of this?  Does separation have to be that complex?  The whole goal here is to have as much as you possibly can in the tool kit to insure you can meet all the capture needs within your organization.  I liken it to buying the a base model with no accessories, and then wishing every day you one or another feature.

So now you have examined your documents, and figured out how to efficiently scan and split.

Wednesday, July 6, 2011

Document Scanning and Capture Planning - Part 1 - Sizing and Storage

Been wanting to do this for quite some time, and finally had some time to sit down and put thoughts together.  I find that many of the scanning and capture implementations lack overall direction, structure and standardization.  I wanted to put together a manual from my experiences, and ask the community to add so we can build a reference for everyone to use.  This will be composed of many parts, including all different topics like storage, hardware, designing your index fields, etc.

Sizing and Storage Planning for Document Management and Scanning



One of the key areas of planning for any scanning/capture implementation is sizing and storage.   Many of the customers we work with have no real grasp on the volume of paper they deal with on a day to day basis, and when they make the migration to digitizing their paper, they are often quite surprised at the amount of paper they push through the system.  Obviously, this can cause some serious issues on many different fronts.   So how do you estimate the amount of paper?  There are several key conversion factors used by the document management industry, as outlined below:

Description
Number of Pages
Storage
1 Scanned Page – 8.5 x 11
1
50KB
1 Scanned Page – 11x17
1
100KB
1 File Cabinet – 4 drawers
10,0000
500MB
1 Box
2500
125MB
1 Linear Inch
100
5MB
1 E Size Engineering Drawing (48x36)
16 – 8.5x11
800KB



This table is a basic planning tool, and can be used as a starting point.  One thing to remember is that these are all standard pages.  Not full image magazine pages, but full text pages.  The other thing to keep in mind is that we have listed for boxes and file cabinets, the average number of pages contained within.  In the imaging world, we deal with images, not pages.  What is the difference?  A page may have 2 sides, which are converted digitally into 2 images.  So effectively, if you have a box with double sided pages you are scanning, you will have to double the storage required.
Some other key factors that can contribute to storage and sizing:

DPI Setting – one of the key questions we always receive is What DPI should I set on my scanner?  For most basic scanning and archive applications, you can set your scanner to 200 DPI.  If you are doing OCR or any type of advanced data extraction, you always want a 300 DPI image for maximum accuracy.  Anything beyond that is just a space killer, will slow down your process and really bloat your files.

Black and White, Greyscale and Color – always use black and white scanning to keep file sizes at an absolute minimum.  Greyscale and color scanning should only be used when absolutely necessary, as file sizes are just crazy.  Below is a table of file sizes for the same letter.  The letter was about 50% page coverage.

Scanning Mode/DPI
File Size
Black and White – 200 DPI
26K
Black and White - 300 DPI
38K
Black and White - 400 DPI
51K
Black and White - 600 DPI
80K
Greyscale – 300 DPI
301K
Color- 300 DPI
577K

Image Processing – image cleanup can significantly reduce file sizes, and it is very important to use this feature whenever you can.  Despeckle, deshade, border removal, etc. will eliminate unnecessary noise in scanned images, and reduce your storage requirement by 10-30% depending on the quality of your documents.

Image Format – There is a lot of misinformation on the market about TIFF versus PDF.  I always hear “We want to store as TIFF because PDFs are just too big.”  Just not the case.  An image scanned to as400 PDF is just a TIFF in PDF clothing (Or a PDF wrapper to be more exact).  The PDF overhead is almost negligible.  The de facto standard in imaging today is rapidly becoming the PDF image with hidden text.  This gives you a nice little file with the pristine image, and converted OCR text in the background.  The text layer adds negligible size to the file.

So now, with all this info, you can estimate volume in images, and then come up with required storage on a monthly, yearly or project basis.

Sunday, July 3, 2011

SharePoint and the Document Management Industry

We are talking denial, and I ain't talking about a river in Egypt (Sorry for the bad joke)

I see it every day, and the misinformation out there about Microsoft SharePoint is just crazy.  First off, let me establish my position.  SharePoint has mapped to the typical Microsoft pattern from a product perspective.  Version 1.0 is usually lacking, causes great pain, and sours many IT folks.  2.0 starts to really get some traction, and people start taking notice, early adopters (also known as gluttons for punishment) go all in, and they continue to gather information for Version 3.  Version 3, they knock it out of the park, address needs, and most IT jump in after service pack 1.  This probably accounts for the slow adoption rates (Some interesting SharePoint Stats here)

SharePoint 2010 and all its wonder is taking business by storm.  It is an incredible tool, when used correctly, as a collaboration tool, document repository and overall business automation tool.  Depending on the business size, structure and industry, what I am finding is that it solves pain points.  Take for example the document capture implementation we just finished.  The customer was looking to eliminate Xerox DocuShare from their organization as they were having too many issues, and could not get adequate support from the vendor.  As a large mining company operating in several large countries in South America, they were having a hell of a time dealing with all their paper invoices, and were looking to automate their scanning and invoice processing. They took a leap, and the pilot project was designed to capture and process invoices within their Chilean AP department.  The project was an absolute success, and implemented within a weeks' time.

Simple.  Effective.  Done.

Now if you were to talk to a traditional Document Management Reseller, or perhaps a vendor, they would have instilled the customer with great fear and doubt:

"SharePoint is not a real document management system."
"The resources required to manage SharePoint will kill any ROI you can glean from automating a process."
"SharePoint cannot handle high volume of documents."

I find the attitude is pervasive, and I think it just comes from a lack of understanding, and truly a lack of effort to research the competition.  Is SharePoint for everyone?  No.  Just as Documentum or FileNet is not for every organization.  But the momentum is absolutely mind numbing.  Watch for profits to fall in the Document Management segment...

Do you think SharePoint is the DM industry killer?

Tuesday, December 28, 2010

Scanning and Capture Models

So, in examining a corporate strategy on how best to deploy a scanning and capture solution, there are typically 3 models:

·         Centralized
·         De-centralized
·         Distributed

Each model has its own pros/cons, and below I will examine each, and dive into some detail.

Centralized
Ah, the centralized model.  Some call this old school scanning and capture, as for many years, this was the only way to get the job done, and convert your paper to digital form.  This model provides a centralized scanning center to provide mass conversion for the organization.  The operation can be run by in house personnel, be managed by a services provider in house, or be outsourced to a scanning service bureau.  It requires high volume/high speed hardware, and typically utilizes advanced capture software to allow for the utmost in automation and efficiency.  The software and hardware operators are typically highly trained, and there are usually only a few of them.  Paper and/or digital media is shipped to the centralized location and processed through a set, standardized capture workflow.

Centralized Pros

·         Easily standardized process due to a limited number of skilled/trained scan operators
·         High speed hardware/software results in minimal processing time once paper is received
·         Centralized reporting and control of overall process
·         No loading on WAN infrastructure
·         Centralized backup and restore

Centralized Cons

·         Usually a high time delay for availability of documents
·         High cost due to shipping of documents
·         High maintenance costs
·         High training costs to bring on new operators
·         Disaster recovery planning issues if centralized site is down
·         Operators are typically not knowledgeable in the documents they are indexing

Decentralized

Over time, as bandwidth and scanning hardware/software prices went down, the obvious move was to decentralize the whole scanning and capture process.  This move placed scanning in the branches, and allowed the whole document capture process to be performed by those who had working knowledge of the documents.  Smaller, desktop class hardware could be used, and most capture companies made batch scanning and upload to the centralized repository simple to accomplish.

Decentralized Pros

·         Scan operators are well versed in the documents they scan
·         Documents are available almost immediately  
·         No shipping or transfer costs for documents
·         Branch control of the whole scanning process

Decentralized Cons

·         Standardization can be an issue
·         No centralized control or reporting
·         WAN Bandwidth consumption can be high
·         Licensing costs can be high depending on software utilized

Distributed

The advance of network-based scanning devices and the lowering of bandwidth pricing led to the newest model, the Distributed Model.  Distributed Scanning allows for just about anyone in the organization to walk up to a network scanning device/scanning copier/fax machine and send documents to a repository.  The devices are typically multi-faceted, and along with repository integration, can provide scan to network folder, FTP and email.  Collaborative back-end systems, like Microsoft SharePoint, lend themselves nicely to this model, as they allow anyone to participate in a Document Workspace.

Distributed Pros

·         Put scanning in the hands of everyone in the organization
·         Provides a great launching pad for collaborative solutions
·         Simple, easy to use interfaces allow for minimal training and quick adoption
·         Capture and indexing is now in the hands of the true document owner
·         One-to-many solution provides a single device to service many users

Distributed Cons

·         Lack of standardization without software addition
·         Security and document control can be major issues
·         Bandwidth from smaller branches can be a problem with larger scans
·         Lack of hardware integrations with back-end systems

So, most organizations today are combining the above models to create a Hybrid Scanning and Capture solution, and leveraging all the strengths together to minimize the weaknesses of any one model.   Another strategy is to tie scanning models to specific business processes, as most lend themselves nicely to specific scanning and capture workflows.
For more information, view a webinar on Distributed Scanning and Capture at the link below: