Part 4: Generative AI in Investigations

Part 4: Generative AI in Investigations

AI Essentials for Investigative Intelligence | 6-Part Video Series

Generative AI is one of the most-hyped capabilities in investigative work, and one of the riskiest to get wrong. Used well, it can compress hours of transcripts, body-worn camera footage, and multilingual content into something teams can actually act on. Used poorly, it fabricates facts and mirrors biases baked into its training data, quietly steering analysts toward incorrect conclusions.

In Part 4 of our AI Essentials for Investigative Intelligence series, we examine generative AI: what it actually does, where it adds value in investigative workflows, and the limitations and risks teams need to account for before deploying it.

Hi, I’m Sean, a Senior AI Product Manager at JSI. Thanks for joining us for Part 4 of this video series, titled AI Essentials for Investigative Intelligence, where we break down the core AI capabilities that matter for investigative work, show where they truly add value, and provide a clear understanding of the considerations that should guide every deployment.

This episode covers the much-hyped topic of generative AI. Generative AI is a sub-branch of artificial intelligence that focuses specifically on creating content rather than only classifying or retrieving it. Depending on the system, that content can include text, images, audio, video, or even software code.

Most modern generative AI systems are built on top of deep learning models, most notably large language models, or LLMs, which I may refer to later in the video. These models are trained on extremely large datasets to learn statistical patterns and relationships in data, such as how language, symbols, or structures tend to co-occur.

In practice, generative AI systems operate in response to prompts. A prompt is simply an instruction or input provided by a user or another system that guides the model on what to generate, how to format it, and what constraints to follow.

From an investigative perspective, it is useful to think of generative AI not as a knowledge source, but as a probabilistic pattern generator that produces outputs based on learned correlations in its training data. In other words, it is not something you should rely on for information and facts on its own, but it can help synthesize large sets of information into something that supports an investigation.

Generative AI supports many different modalities, or data types, for both inputs and outputs. The most common form factor is text. A classic example is ChatGPT, where users, or in some cases systems, interact with the model through a text input, also called a prompt. For example, I can provide a prompt asking the model to write a brief overview of this video series, and it will generate a description based on the context and constraints I provided.

Generative AI models are not limited to text inputs and outputs. Other data types are also relevant for technical investigators to understand, starting with voice and audio content. For example, platforms such as ElevenLabs can be used to create voice clones through text-to-speech models. Once an initial voice model exists, these systems can generate outputs in different languages as well.

This demonstrates how the technology could be abused and why it may be relevant to financial fraud or other fraudulent activity. It is an important capability for technical investigators to be aware of, especially as synthetic voice content becomes easier to create and harder to distinguish from authentic recordings.

Image content can also be used both as an input for visual analysis and as an output for creating new content. This is especially an area of concern for law enforcement as generated image content becomes increasingly difficult to distinguish from reality. In the example shown, all of the images were created using a generative AI model and were not real.

Within the last year, we have also seen a significant improvement in the quality of video content produced by generative AI models. This is important for investigators because synthetic video can introduce noise or misleading narratives into an investigation. For example, one video shown was fully produced by Google’s Veo 3 model using only a short text prompt as input. There were no real actors or voices used; the entire fabricated car show was generated by the model.

Generative AI models are also being used for coding, which can be useful for technical investigators looking to work with large datasets or automate analytical processes. At the same time, the industry is moving toward multimodal models, which can handle multiple data types as both inputs and outputs. For example, a multimodal model could take an image of a real person as input and generate a video of that person doing something described in a text prompt. These models are becoming more general in the types of data they can interpret and the types of content they can generate.

When applying generative AI tools in an investigative context, much of the focus today is on using large language models to support key capabilities. These include summarization, where LLMs can condense long transcripts, interviews, and body camera footage into key points so investigators do not need to manually review all content to find segments of interest.

They can also support classification tasks, such as automatically categorizing records, highlighting high-priority items, and detecting sentiment or emotion. These capabilities can help investigators focus on what matters within a larger dataset.

Large language models are also becoming effective at translation tasks, including translating messy or foreign-language text into something understandable. This can help bridge language barriers that investigators may encounter in their daily work.

Finally, these models can support analysis by helping generate concise search results, extract key entities or topics, and create on-demand visualizations that turn raw data into actionable insights. These are only a few examples of the use cases currently appearing in investigative products, and we can expect a significant increase in the applications enabled by this technology.

Of course, generative AI is not without challenges. One major challenge is accuracy and hallucinations. This was a significant issue for some early products that used large language models behind the scenes. In some cases, these systems would completely fabricate facts or create references to articles or journals that never existed.

In the next video, we will discuss techniques and design patterns that can help mitigate or account for these limitations. However, it is important to be aware of these risks if you are planning to use generative AI tools to answer questions or act as a knowledge base, which is not fully advisable without appropriate safeguards.

Another major challenge in any AI deployment is bias, fairness, and censorship. For example, some large language models developed by Chinese companies may be designed not to criticize the ruling government. This is an example of guardrails imposed directly within the model that users cannot easily override.

Because of this, it is important to understand what data was used to train the model, what biases may be inherent in that training data, and how diverse the dataset was. These are all important questions when considering the use of generative AI within an operational environment.

There are also issues around copyright and intellectual property associated with the underlying training data. For example, lawsuits have emerged around the use of copyrighted news articles in AI training datasets. As a result, transparency around which models are being used and how they were trained is an important consideration when rolling out these tools in operational environments.

Finally, the rise of deepfakes and other synthetic media will become increasingly problematic as it becomes easier to create large volumes of misleading, unsavoury, or graphic content. This can create challenges around scale, because investigators need to focus on data involving real victims who need help, rather than fabricated content. Sorting through that noise and volume of data will likely become more difficult before it improves, so it is important to approach this technology with eyes wide open.

When using tools built with generative AI in high-risk technical investigations, several additional requirements should always be considered. First, responses from these tools should be grounded in collected data, and users should be able to trace any given fact back to the source data. Human users should have the final say, rather than blindly accepting generated output as fact.

Inline citations or links back to source data are therefore incredibly important in these systems. It is also important to ensure that the underlying LLMs do not get in the way of an investigation. Investigators may be dealing with highly sensitive content, so the models being used should not prohibit interaction with that data because of inherent censorship or guardrails.

Lastly, all access to underlying data and information must fully respect the access control policies configured within the organization. Complete and exhaustive logging is also essential so data protection officers or administrators can audit and review the effective use of these tools.

That wraps up Part 4 on generative AI. So far in this series, we have discussed the data problem that law enforcement and intelligence agencies are up against, where AI fits into that challenge, and the AI subfields of natural language processing, computer vision, and generative AI.

If you have not seen the earlier episodes yet, I recommend watching them before moving into the next episode on emerging AI systems. That is where we will show how these different AI capabilities come together into real workflows and what that means for operational impact. We will conclude with Part 6, which will review compliance and ethics.

Thank you for watching. I hope to see you in the next part. If you want to learn more about JSI, please visit jsitelecom.com.

What You’ll Learn in Part 4

  • What generative AI actually is, why it should be treated as a probabilistic pattern generator rather than a source of truth, and how that mindset shapes responsible use
  • The core generative AI modalities investigators need to know (text, voice, image, video, and code) and how they combine into multimodal systems transforming investigative workflows
  • Where LLMs deliver value in investigations, including summarization of long transcripts and footage, classification of priority items, translation, and generating new content
  • The risks and safeguards that determine whether generative AI is deployment-ready, including hallucinations, bias and censorship, source grounding with inline citations, and audit-ready access controls

What’s Ahead in This Series
This is part 4 of our 6-part series on AI Essentials for Investigative Intelligence.

Already published:

Coming next:

  • Emerging AI Systems: How multiple AI capabilities integrate within real investigative workflows to deliver operational impact.
  • Compliance and Ethics: Why responsible, transparent, and policy-aligned use of AI is mission-critical in high-risk environments.
Sean Thibert
Sean Thibert Senior AI Product Manager

Sean Thibert leads JSI’s AI product strategy and execution, enabling customers to streamline workflows and uncover insights from complex datasets using advanced technologies. Over the last five years, Sean has worked closely with public safety agencies around the world to deliver secure, compliant digital intelligence solutions. Prior to joining JSI, he developed automated data pipelines for geospatial and imagery analysis, integrating machine learning models to accelerate processing and improve accuracy.

You might also like

Investigative intelligence feature image for part 1 on the data problem law enforcement and intelligence teams are up against.

Part 1: How AI Can Help Reduce Data Complexity in Investigations

AI Essentials for Investigative Intelligence | 6-Part Video Series Law enforcement and intelligence teams are drowning in data. Today’s investigations include more sources, more(...)

Read more
financial crime on-demand demo landing page thumbnail image, JSI

Following the Money: Exposing Criminal Networks with 4Sight

Watch how analysts use 4Sight to follow the money, connecting fragmented data and identifying a criminal network.(...)

Read more

AI for Law Enforcement and Intelligence: The Critical First Step Most Leaders Miss

Skeptical that AI can make your operations more efficient? That's a good sign.(...)

Read more