Blog

Software tips, techniques, and news.

FileMaker Amazon Textract Integration

Amazon Textract is a service that automatically extracts text and data from scanned documents, going beyond simple optical character recognition (OCR) to identify, understand, and extract data from forms and tables. Many companies today collect data from scanned documents, such as PDFs, tables, and forms, via manual data entry, which is slow, expensive, and error-prone. The Textract service makes it easy for customers to accurately process millions of document pages in just a few hours, significantly reducing document processing costs and enabling them to focus on deriving business value from their text and data rather than wasting time and effort on post-processing. Integrating your FileMaker application with Textract lets you quickly and easily leverage AWS machine learning to convert paper or scanned documents into useful digital information.

youtube-preview

How to Get Started

In order to begin connecting your FileMaker database to work with AWS Textract, you will first need to have an active AWS account and set up an IAM user. Amazon has a great walkthrough on setting up these profiles, and be sure to capture the Secret Key and Access Key for your user.

AWS IAM User Example.

You will also want to ensure you have granted full access permissions for Amazon S3, where our processing documents will be stored, as well as for Textract. Once your user is created, you will also want to navigate to the S3 service and create a new storage bucket that your user has full access to. Capture the name and region of your new bucket, and store them along with your users' keys in your FileMaker database to connect. Please note that Textract does not need to be explicitly enabled on your AWS account and has no minimum fees, but it is a paid service that charges around 5 cents per processed page. There are free-tier options that may apply to your account, so feel free to explore them in your AWS management console. 

Uploading your Document

Once you have the information from AWS, choose a document to process and store it in a container field in your FileMaker database. You will want to use an image or a PDF, and I would suggest starting with a simple table or form to understand how data is processed and returned by Textract.

To process your document, you need to store it in the S3 bucket we created earlier.

amazon s3 bucket example.

This is done by using cURL to send a hash of the data, along with a signature that includes your user keys, to the Amazon S3 host, and specifying the bucket and region to store it in. Note that there is no session token authentication process, as you might see with many other integrations; instead, each request delivered to AWS includes a signature containing your specific user key information. A great example of building AWS cURL requests, widely used throughout the FileMaker community, is included in the demo you can download below.

Running the Textract Analysis

Now that your document has been uploaded and stored in an S3 bucket, the next step is simply telling AWS to trigger the Textract Document Analysis job. This is done again by sending a cURL request with your user signature to the AWS host, but this time specifying the Textract service and pointing to the document you want processed. Once the job has been successfully initiated, AWS will return a Job ID that you should store as a reference for the Textract process you just started and the data it will produce. Depending on the size and complexity of the document you are uploading, the analysis process can take anywhere from a second or two to well over a minute.

textract data return.

There are a couple of ways to determine when the process is complete, but the simplest is to request a status update from AWS. By providing the Job ID in another cURL request to the Textract host, you can check whether the job status is "IN_PROGRESS" or "SUCCEEDED". Once the analysis is complete and the job status is successful, the response will also include "BLOCK" data. This is the information we are after: a large volume of JSON data indicating the type, position, and content of each recognized element in our document, along with its relationships to other elements.

Processing your Block Data

The Textract process will generally return information about your document in three different types: raw text, table content, and form data (also known as key-value pairs). The best way to turn this into something we can use in FileMaker is to loop through the large volume of JSON and create a record in a separate BLOCK table for each element. Here you can store content, type, location, and relationship information for each element, and then we can use simple relationships to link parent and child elements. This will provide a much more useful way to either present analyzed information to a user or parse the data we care about into the relevant FileMaker fields.

processed textract data.

Conclusion

Integrating Amazon's Textract service with FileMaker can be an excellent, low-cost way to go beyond 3rd-party OCR and pull relevant information from images or PDFs directly into your solution. With all processing done automatically, you can eliminate double data entry and produce useful reports and analyses much more quickly.

Ready to put Textract to work in your FileMaker solution? Here's how we make that happen. Contact us if you would like help integrating Amazon Textract into your FileMaker application!

Did you know we are an authorized reseller for Claris FileMaker Licensing?
Contact us to discuss upgrading your Claris FileMaker software.

Download the DB Services - FileMaker Amazon Textract Integration File

Please complete the form below to download your FREE FileMaker file.

First Name *
Last Name *
Company
Email *
Phone *
FileMaker Experience *
Agree to Terms *
brandon terrell headshot.
Brandon Terrell

Brandon is an energetic FileMaker developer with a natural ability to connect with people.