The entire document lifecycle, in a single platform.
From scan to publish: Captika captures from any source, extracts and validates data with OCR and AI, recognizes the document type and publishes the result wherever you need it. With rules, .NET scripting and end-to-end traceability.
Capture from any source
Connect Captika to the channels you already use and automate document intake without changing how you work. One flow, many entry points.
- Physical TWAIN scanners with real-time scanning and image quality control.
- Office 365 email: process attachments directly from corporate mailboxes in the cloud.
- Local or network File System with folders and subfolders, in-use file control and automatic renaming.
- FTP / SFTP servers, ideal for automating distributed environments.
- SharePoint via ViewXML searches and Thuban from work trays, queries and Stored Procedures.
Capture sources
Set up one or several sources per capture job and let them run automatically.
Scripting engine
Customize every event of the process and connect to any external system. Virtually unlimited integration possibilities.
Correct data from the start
Define metadata with full flexibility and validate information against your systems before publishing it. Fewer errors, less rework.
- Three field levels: by batch, by document and by page.
- Smart drop-down lists generated from databases (dynamic SQL), Thuban or SharePoint lists.
- Automatic validation and enrichment with SQL queries, revalidation against Thuban or custom scripting.
- .NET 10 scripting engine (C# / VB.NET) for rules and real-time calls to WebServices or APIs.
Top-tier OCR and AI
Extract key information from any document with the most powerful OCR engines on the market, and pick the best one for each document type.
Tesseract 5x (native)
Local OCR with spatial capture of every word and line, avoiding redundant reads and enabling on-screen visual validation.
Amazon Textract
Cloud OCR for printed and handwritten text, reading tables, forms, signatures, faces and labels in images.
Google Vision AI
Recognition via API, natively integrated, to add another leading market engine to your flow.
OpenAI / Bedrock
Run smart prompts over the extracted text to improve data extraction and recognition.
PDFs with text layer
With PDFBox and PDFIUM: split pages, detect signatures and leverage the original text layer for more speed and quality.
Barcodes
1D, 2D and PatchCodes reading: CODE128, CODE39, DATAMATRIX, PDF417, EAN13, QR and more, to split and classify automatically.
Dynamic templates for any document
Captika recognizes both structured documents (invoices, tax forms, promissory notes) and unstructured ones (minutes, balance sheets, letters, contracts), with extraction rules per field.
- Identify documents by barcode, extracted fields, page order, size, weight or shape (image fingerprint).
- Text similarity: upload a sample image and set the minimum similarity level to recognize future documents.
- Word bags per document section, to dramatically speed up template building.
- NLP cleaning rules: dates, emails, spelled-out numbers, tax IDs (CUIT / CUIL / RUT), data masks and more.
Applied AI (Amazon + OpenAI)
Combining these tools, your capture pipeline learns and improves over time.
Publish to multiple destinations, in parallel
Generate multi-page PDF or TIF —with a text layer created at capture time— and publish documents and metadata simultaneously to several systems, each with its own configuration.
FileSystem
Dynamic folder structures based on recognized fields, with configurable file names and parameterizable text files.
SQL databases
Publish all captured data for use in analytics, management or back-office systems.
FTP / SFTP
Secure transfer of documents and data, with fully configurable names and exported metadata.
Thuban Portal
Direct indexing: upload new documents or update existing records.
Microsoft SharePoint
Native mapping of Captika fields to the metadata configured in your SharePoint lists.
Amazon S3
Cloud publishing by selecting the destination S3 bucket and the file name.
Captika Desktop & Captika Service
Human-assisted capture and unattended automation. Use them separately or combine them into a hybrid solution.
Assisted capture
For operators who validate documents visually: digitization centers, administrative areas and branch offices.
- Real-time scanning from TWAIN scanners.
- Visual validation of images and manual correction of OCR data.
- Quality control: blank pages, rotation and resolution.
- Fields by document or page, multi-user with profiles and traceability.
24/7 automation
A Windows service to process large volumes 100% automatically, without human intervention.
- Multithread processing of multiple jobs at once.
- Automatic capture from folders, Thuban and Microsoft 365 email.
- Task schedule: define what runs and when.
- Error handling, retries and complete logs for auditing.
| Feature | Desktop | Service |
|---|---|---|
| Built-in OCR (Tesseract) | ✓ | ✓ |
| OCR/ICR via API (Google Vision, Amazon Textract) | ✓ | ✓ |
| C# / .NET 10 scripting engine | ✓ | ✓ |
| External validations (DB / API) | ✓ | ✓ |
| Detailed audit logs | ✓ | ✓ |
| User profiles and roles | ✓ | ✓ |
| Automatic template updates | ✓ | ✓ |
| Visual review interfaces | ✓ | — |
| Horizontal scalability (multi-instance) | — | ✓ |
| 24/7 execution, no intervention | — | ✓ |
Cloud engines (Google Vision, Amazon Textract, OpenAI / Bedrock) run on the client's own credentials and their usage is billed directly by the provider, with no intermediation or markup from Vivatia. Local OCR with Tesseract has no per-page cost.
Choose Desktop if you need…
- Human validation and real-time data correction.
- A friendly visual interface for non-technical users.
Choose Service if you want…
- Automated, continuous processing at scale.
- Unattended integration with multiple sources and destinations.
Try Captika with your own documents
Start free with the Free Edition or request a quote for Captika Server for bulk processing.
