In an increasingly digital world, video calls are becoming the norm in everyday life. However, extracting important information from these meetings can be a tedious task, often requiring the review of recordings or note-taking during the meeting. Even so, recording the key points discussed and the next steps defined is crucial for the continuity of projects and organizational efficiency, making it a mandatory step.
With the advancement of AI technologies, specifically Large Language Models (LLM) such as GPT-4, new possibilities have emerged in the handling and extraction of information from unstructured data. Such tools allow the creation of innovative solutions for complex problems, where conventional methods and technologies cannot be effective.
Given this scenario, the innovations brought by Large Language Models present themselves as valuable tools. They enable the transformation of unstructured data generated in video calls into actionable insights and organized records, thus paving the way for more effective and automated information management.
Introduction to the Solution
The proposed solution to the problem uses a serverless architecture with AWS services. The solution follows this flow:
When a video is uploaded to Amazon S3, a Lambda is triggered by the Bucket’s Event Notification
The function extracts the audio from the video and saves the file in another folder of the Bucket.
This file triggers an Event Notification for another Lambda.
The Lambda creates a job in Amazon Transcribe.
The result of the job is sent to the Bucket.
A Lambda is triggered by the job result, it applies some processing and sends the text with a prompt to the Amazon Bedrock API, the prompt is written with the goal of extracting the main points discussed in the meeting, and if there are any, what the next steps defined were.
Extracting the Audio
The first step to automating the extraction of information from a video call is to obtain the audio from the video. This process starts with an Event Notification, so that when a video is uploaded to the S3 Bucket, a Lambda function is triggered.
The Lambda function then uses a Python library to extract the audio from the video. Once extracted, the audio is stored in the same S3 Bucket, but in a different folder, in order to keep a clear organization of the data.
This new audio file triggers the next stage of the process: the execution of a job in Amazon Transcribe. This is where the audio will be transcribed to text, enabling the extraction and analysis of the information from a video call by an LLM.
Running a Job in Transcribe
Amazon Transcribe is a powerful speech-to-text conversion tool, designed to simplify the audio transcription process. This service uses the concept of jobs, which are specific configurations for processing the desired audio files.
When creating a job in Amazon Transcribe, the user needs to provide essential information, such as the location of the audio file, the language spoken in the audio, and, if necessary, can also use a custom vocabulary to ensure the accuracy of the transcription, especially in cases involving technical terms or specific jargon.
The transcription process itself starts when the job is triggered. At that moment, the audio file is submitted to the system, which uses advanced speech recognition algorithms to convert speech into text. When the job is finished, the result is a file in JSON format that contains not only the textual transcription, but also additional information such as the confidence in the accuracy of the transcription (indicating how reliable the transcribed text is) and the timestamp of each word.
After the transcription process is complete, the next step is the use of Amazon Bedrock. Amazon Bedrock plays an essential role in the continuation of the process. Once the text has been transcribed, it is time to extract insights and perform deeper analyses.
Using Amazon Bedrock
With the text transcribed by Amazon Transcribe in hand, the flow continues with a new Event Notification that activates a Lambda, which makes a call to the Amazon Bedrock API. This AWS managed service makes it easier to build and scale generative AI applications using foundation models. It offers several model choices from various leading AI organizations, allowing updates with few code changes.
Amazon Bedrock allows easy customization of models with your own data without the need for programming. You can build managed agents that invoke APIs to perform complex tasks, and it also has native support for Retrieval Augmented Generation (RAG) to connect models to proprietary databases, ensuring data security and compliance with standards such as HIPAA and GDPR.
The model chosen for the solution is Claude 2.0, which is a large-scale language model, with an impressive context window of 100K tokens, which allows a detailed analysis of large blocks of text. Its ability to process and understand complex contexts makes it a valuable tool in extracting critical information from the transcribed text.
At the core of Bedrock, Claude 2.0 is triggered with a predefined prompt, which guides the model on what to look for in the transcribed text. This prompt can be configured to identify and extract key points discussed and the future actions defined during the meeting, providing a clear and concise view of the content discussed.
Large Language Models (LLM) such as Claude 2.0 are known for their versatility and ability to adapt to a variety of natural language processing tasks. This flexibility makes Amazon Bedrock a robust and adaptable platform to handle the challenges of text analysis, making it possible to transform meeting transcriptions into actionable insights effectively and efficiently.
By using Amazon Bedrock, organizations not only manage to automate the extraction of information from meetings, but also expand their text analysis capabilities, enabling a deeper and more contextualized analysis of unstructured data.
Conclusion
The proposed solution offers a fully serverless system that relieves a significant business pain, automating the extraction of information from recorded meetings. In addition, it integrates seamlessly with AWS services, providing a simplified implementation and efficient operation. By leveraging the power of Amazon Bedrock and Large Language Models, organizations can now transform their video conferences into actionable insights, thus boosting productivity and informed decision-making.

