When your doc repository incorporates tons of of thousands and thousands of recordsdata amassed over practically a decade, how do you systematically discover and redact delicate buyer knowledge with out taking years to finish? This was the problem dealing with The Huntington Nationwide Financial institution (Huntington), a prime 10 financial institution in the US.
Redacting delicate data at scale
Since 2015, Huntington’s doc administration system has securely saved tons of of thousands and thousands of paperwork on-premises. In 2025, as a part of a proactive compliance initiative, Huntington got down to course of the paperwork on this system and redact delicate knowledge. These paperwork come in numerous codecs, so the answer wanted flexibility to deal with assorted file sorts whereas delivering the throughput required to course of thousands and thousands of paperwork rapidly.
Unique estimates indicated this effort would take years. Nonetheless, by designing a scalable redaction workflow utilizing Amazon Textract, Amazon SageMaker, AWS Step Capabilities, and AWS Lambda, Huntington diminished this timeline to months.
Resolution overview
Earlier than inspecting the technical implementation, let’s have a look at the core necessities Huntington established for this venture. When you’re dealing with an analogous large-scale doc processing problem, these necessities can function a place to begin to your personal answer design:
Knowledge have to be encrypted at relaxation and in transit.
Areas the place knowledge is saved or accessed should meet strict entry necessities.
Companies used have to be in-scope for PCI DSS compliance.
Outputs have to be replicated again to on-premises knowledge shops.
Redaction accuracy should meet or exceed 95% to satisfy compliance necessities.
The next diagram illustrates the high-level answer structure.

Shifting knowledge securely, with confidence
Huntington’s first goal was to maneuver paperwork from an on-premises file share to an Amazon Easy Storage Service (Amazon S3) bucket. Shifting paperwork is easy, however this effort required transferring over 400 million paperwork, encrypted in transit and at relaxation. To perform this, Huntington used AWS DataSync, AWS Direct Join, Amazon S3, and AWS Key Administration Service (AWS KMS).
AWS DataSync may be deployed as an agent in your on-premises knowledge heart to watch a configured supply, corresponding to an SMB file share. Whereas getting paperwork to AWS was vital for processing, AWS DataSync additionally helps syncing knowledge again to on-premises, which was one other key requirement for this venture.

Amazon Textract is an AWS machine studying service that extracts textual content, tables, and varieties from scanned paperwork. Monetary establishments use it to mechanically course of paperwork like account statements or mortgage purposes, then establish delicate knowledge corresponding to Social Safety numbers, account numbers, and private addresses. The next pattern bill demonstrates this functionality.


Amazon Textract detects varied fields from a doc and supplies coordinates of detected fields and different metadata inside a JSON output.
Huntington used Amazon Textract in an orchestrated course of with AWS Step Capabilities. This strategy diminished guide overview time whereas bettering accuracy in detecting delicate data throughout massive doc volumes.
Scaling detection throughput
Automated pipelines for doc processing are worthwhile, however processing paperwork sequentially would have prolonged the venture timeline to years. To satisfy their objective, Huntington wanted to course of thousands and thousands of paperwork every day.
Scaling to this stage required addressing two essential concerns: maximizing concurrent Amazon Textract jobs inside service quotas, and controlling request charges to keep away from throttling.
AWS companies have quotas that may be adjusted by means of mushy and laborious limits. The Amazon Textract jobs-per-second quota may be elevated by submitting a request by means of the AWS Service Quotas console.
To maximise throughput in opposition to the service quota, Huntington used the AWS Step Capabilities built-in map state, which processes collections of inputs in JSON, CSV, or different codecs. The staff organized paperwork in Amazon S3 right into a JSON assortment and ran the map state in distributed mode for greater concurrency. To trace pipeline progress, they used AWS Step Capabilities map run execution summaries alongside Amazon CloudWatch dashboards to watch response occasions, throttle counts, successes, and error charges.
To deal with potential throttling, Huntington monitored their CloudWatch dashboard to confirm Amazon Textract profitable request counts and throttled counts. As wanted, they adjusted concurrency limits for youngster workflow executions to substantiate they remained below the Amazon Textract service quota whereas sustaining excessive throughput. When jobs accomplished efficiently, detected fields and metadata have been written to a bucket for later overview. The next diagram depicts this strategy:

The wait block inside the step operate verified the method was able to proceed with writing job metadata and persevering with with the following Amazon Textract invocation. When there are not any failures, the state machine finishes with a move state. When failures happen, AWS Step Capabilities writes to a log for human overview and reprocessing.
Redacting detected delicate data
Up so far, the method centered on detecting delicate knowledge and cataloging it inside metadata recordsdata written to Amazon S3. The ultimate steps are to redact the paperwork and transmit them again to on-premises storage.
Picture and PDF redaction is supported by a number of open-source and proprietary instruments. Widespread open-source Python libraries embrace PyMuPDF or picture drawing libraries like PIL. The next determine exhibits a pattern redaction of the bill proven earlier. Amazon Textract helps detection of assorted fields, and you too can create customized classifications utilizing regex patterns. Mixed with redaction software program, you may confidently redact detected fields. If you wish to create a threshold for human intervention, Amazon Textract supplies confidence scores that may set off validation workflows.

As soon as once more, Huntington confronted the identical architectural problem: how would this scale? AWS Step Capabilities offered the answer for processing thousands and thousands of paperwork whereas providing hooks for error dealing with and retry logic. Because the doc processing pipeline cataloged objects requiring redaction, Huntington ran a easy stream in opposition to them:

To confirm accuracy and thoroughness, Huntington double-checked that detected fields matched anticipated patterns previous to redaction, adopted by a metadata replace for every file. Redacted recordsdata have been positioned in an Amazon S3 location monitored by AWS DataSync for transmission again to on-premises file storage.
Conclusion
Utilizing AWS, Huntington processed paperwork at a price of roughly 10 million per day, lowering estimated processing time from years to only a few months. The price of processing your entire doc repository was roughly 5% of the unique estimate. Redaction accuracy exceeded 95%, assembly compliance necessities and supporting knowledge safety targets.
This venture demonstrates how AWS companies can assist large-scale knowledge processing and compliance initiatives. Huntington plans to proceed utilizing this framework for high-volume redaction wants corresponding to mergers and acquisitions.
To be taught extra concerning the companies used on this answer, go to the Amazon Textract element web page or discover the AWS Step Capabilities documentation.
Acknowledgements
Particular due to the next people and groups for his or her contributions: Xuelei Yuan, Robert Carnell, Jeanne Keith, Debbie Montgomery, Invoice Gross, Jodi Pettiford, Jon Glazer, Marshall Doss, Bob Wojasinski, Tami Wolf, Marijane Eldridge, Pradeep Kumar Tata, Michael Burkhardt, Nirmal Antony, Trevor Pease, Bryan Griffith, Angus Ferguson (AWS) Randy Patrick (AWS), Stephanie Brenneman (AWS), Artwork Steele, Kevin Owen.








