Theta Lake’s ML playbook for multimodal compliance

ML

RegTech firm Theta Lake has set out how its machine learning architecture has evolved to police an increasingly multimodal communications landscape, arguing that genuinely AI-native platforms are pulling away from rivals that bolt AI onto legacy systems.

The company positions artificial intelligence as its foundational infrastructure rather than an add-on.

Theta Lake recently discussed the evolution of machine learning at the company and building a multimodal foundation.

Its first corporate hire was a chief data scientist, its compliance classifiers have used AI since inception, and its architecture rests on proprietary patents dating back to 2018. The firm has been named a Visionary in the Gartner Magic Quadrant and holds independently verified ISO 42001 and CSA Star for AI Level 2 certifications.

According to Theta Lake distinguished engineer Rohit Jain, the platform was designed from the outset for a multimodal world.

Legacy lexicon-based systems built for email, the firm argues, cannot cope with the nuance of modern channels such as video and chat. Because they cannot generalise beyond syntax into semantics, they generate very high false positive rates, lacking any awareness of context.

Theta Lake’s first major engineering hurdle was moving past rigid keyword detection. Rather than forcing customers to abandon their existing lexicons, the company developed proprietary IP (U.S. Patent No. 12,045,561) to handle lookalikes and soundalikes, keeping the system robust against misspellings, OCR errors and semantic variations. Further IP followed for fuzzy matching across larger text sections, along with error models for OCR and transcription (U.S. Patent No. 12,265,563), supporting use cases such as disclaimer detection and transcript matching.

Recognising that model quality depends on data quality, the firm rigorously benchmarked third-party OCR and transcription engines across varied audio and video conditions, speakers and accents, building a bespoke transcription test suite it runs periodically to select the best-performing vendors.

On the modelling side, Theta Lake progressed from word embeddings to sentence embeddings, selecting those with the optimal balance of sensitivity and specificity. These embeddings feed ensembles spanning Naive Bayes, neural networks, tree-based methods, boosting, GBMs and KNeighbors, in a process that has moved from manual to almost fully automated.

Where accuracy gains justify the performance overhead, fine-tuned discriminative and generative language models are deployed selectively. Every classifier is now a custom-tuned ensemble, refined through a data- and metrics-driven methodology. Similar principles govern its visual pipeline, from object detection to image classification, including patented methods for detecting applications shared on screen (U.S. Patent No. 12,464,032 B2).

Read the full Theta Lake post here. 

Read the daily RegTech news

Copyright © 2026 RegTech Analyst

Enjoyed the story? 

Subscribe to our weekly RegTech newsletter and get the latest industry news & research

Copyright © 2018 RegTech Analyst

Investors

The following investor(s) were tagged in this article.