Moderator's introduction
Could I just have the clicker back. And our next topic is the closing one for today, but it's more relevant than ever, because, I think, last year everyone and their dog rolled out artificial intelligence. Today, basically, that has brought a lot of useful things, a lot of interesting things, but also all kinds of challenges and threats. Which is what we'll hear about today from a representative of conference partner STC, Mikhail Inkin. Please, Mikhail, the floor is yours. Let's give him a round of applause. The bottom clicker.
Talk and Q&A
Good afternoon, everyone. Yes, hello. We've got technical problems here. Can we go to the first slide?
Let me flip through it myself, we've jumped somewhere here. Well, now you've seen all my slides at once, that's not bad either. A teaser. Yes, a little teaser. Somehow we started from the end, not the beginning.
Thanks. Good afternoon once again. My name is Mikhail, STC, head of a project group. And today I'll be glad to tell you how we use our biometric technology to help solve forensic tasks. A couple of words about STC. We've been on the market for over 35 years. We make products and solutions based on voice biometrics, face recognition technology and intelligent speech technologies. We've delivered over 5,000 projects worldwide, including the Russian Federation and many other countries. I'd note separately that our algorithms today are internationally recognised, they regularly take part in international scientific and technology competitions and prove their effectiveness on the international stage. Some of those competitions are listed here on the slide. NIST, the CHiME Challenge.
We have fairly broad expertise in working with large language models, including Sber's GigaChat model and also freely available open large language models. On top of that, we have expertise in fine-tuning large language models to a specific customer's request, to their practical, applied task. And today in my talk I wanted to focus on how audio data can enrich an expert's work, what additional insights this audio data can give experts. And I'll probably start with the question of voice biometrics in general. And briefly, it should be said that, overall, voice biometrics can be divided into several main methods. There are automatic recognition methods, where the work is done in automatic mode with minimal operator involvement.
Today we produce systems that are language-independent and text-independent. Whatever language the speaker speaks, our biometrics will recognise them. And, of course, there's the principle that the conclusions of any automatic system need expert confirmation. Of course, following that principle, we also offer our customers both automated speaker identification methods and fully manual methods. They're implemented in our products. Examples of the technologies we use in our automatic algorithms. And it must be said, our technologies work on different types of microphones and perform quite well in different acoustic conditions.
Just a couple of words about speech and the origins of biometrics in general. It should be noted that more than 70 parts of the human body are involved in producing the voice. Every person's body is unique. Because of that, the voice each person produces is unique, which makes it possible to tell speakers' recordings apart by automatic methods, and by manual methods. For example, for more than 30 years, audio recordings, forensic speaker examinations have been used as forensic evidence both in Russia and abroad. I'll probably skip the question of voice vectorisation, skip the question of how automatic models work. I'll say a bit more about manual approaches to speaker identification. This is mostly a forensic audience here. Clearly, a voice has different representations. The voice, an oscillogram, a spectrogram, a cepstrogram – these are all different representations of the same recording of a human voice. And it should be noted that a voice has identifying characteristics, for example, such as the formants, which are examined, for example, the fundamental frequency, which, when you work with them, let you reach unambiguous conclusions in the course of a forensic speaker examination.
We also can't ignore the question of how difficult this task is, because, as one of the biometric modalities that is easiest to obtain, a voice is easy to record: set down a recorder, a mic, grab a phone recording. At the same time, this speaker identification task is quite difficult and it's complicated by a large group of factors. These are speaker factors themselves, technological, say, different microphones have different frequency responses. Communicative ones, well, in one case I'm talking to my child, in another with my accomplice, in a third with a gang member, say, if I'm a criminal. And my voice will sound completely different in each of those communicative situations. Plus there's the group of technical factors: SNR, presence of noise, the speaker-to-microphone distance. All of this brings great variety into the audio data our users may have. And, of course, let's note that today our algorithms can overcome these difficulties quite successfully and perform identification even in difficult acoustic conditions.
Our products use the following methods, the main ones for identification. The auditory method, listening; acoustic; the phonetic method of analysing the fundamental frequency contour; the spectrographic and automatic methods. And today I want to talk about how different STC products solve forensic tasks. Probably, on some practical case, you can imagine that there is a set of audio data. It could have been obtained in different ways. Maybe it's lawful interception of phone calls, maybe it's microphone recordings, maybe it's a seized device and voice messages were pulled from it. There's some body of data, and we want to analyse it. It contains various unprocessed voices, some dialogues, some voice messages.
What do our solutions let you do with these data sets? First of all, let's say that we offer products for all the key stages of the forensic process. Today the main focus will be on data analysis with the AVIS product, and on obtaining evidence. That's the IKAR Lab 3 forensic suites, and also on the Nestor AI solution. We won't touch on data collection now, because that's a separate topic, and one could talk about it at length too. Let's start with the question of data analysis. These data sets can be analysed, for example, on the AVIS system. That's our search and analytics system, which in automatic mode, from raw, unprocessed audio data can extract a large amount of information. What can you do, for example? For example, you can automatically get information about the speaker's gender, the conversation language.
You can get the recognised text. Today we support 18 languages. Russian, naturally, the CIS languages: Ukrainian, Kazakh and some others. And also, for example, foreign ones, Arabic. You can run a biometric search, translate text from a foreign language into Russian, and once you've got the necessary data, you can go on to run analysis functions, for example, cluster analysis, statistical analysis, and do link analysis. On the following slides I'll show a few examples of how such a system works. Here is the application's working screen, and it shows the process of how a user of the system filters a large amount of data.
The system finds only those recordings that match the search criteria. For example, recordings that contain a particular voice, that contain particular keywords. Here they're highlighted on the oscillogram. There's a message transcript. You can run link analysis on all these recordings. For example, produce analytical graphs like these. And as one example of an analytical graph, it's possible to show a link between topics, what topics the voices of interest to us were talking about. Here one such recording is highlighted. It's definitely not her. His second one was a young junkie girl.
Definitely not her. Here's an example of that recording. Maybe we could turn the volume up a bit for the next examples. Here you can see that the keyword was triggered. And, accordingly, we need to go back.
Thank you. And this keyword that was found made it possible to assign this conversation to some specific topic of interest. Let's skip the next few pictures. Next, I want to talk about how large language models help improve search and analytics systems like these, the ones based on voice biometrics. The main capabilities of LLMs are probably already widely known. I won't dwell on this slide and will show a few practical cases of applying large language models in biometric search and analytics systems like this.
One approach is plugging in LLMs, large language models, to extract the information of interest from a body of data. That is, using a specific prompt, the user can analyze the data set that was found and get a brief summary of all the recordings found, highlighting only the recordings that are of the greatest interest to them. So the user reads a short summary instead of listening to every recording and reading the contents of each one, understands which of these recordings are of the greatest interest to them, and then works with each individual recording for further examination. Another use case, it's also shown on the screen here right now, is summarizing the text of a conversation. The process is as follows. Text is obtained from the audio, then the text is fed into the language model, and, via a prompt, the result comes back as a condensed textual representation.
I also want to show separately on this slide, for example, that our experiments with large language models show that large language models are quite good at extracting the information that is of interest to us. As one example, I suggest looking at item number 5 here. Code words and slang used. So, for example, large language models today are quite good at detecting recordings where some kind of concealed language is used, where certain objects or items are being disguised. And this gives new insights to the users of such systems.
Of course, I should note here that search and analytics systems solve search tasks, whereas evidentiary tasks, for those we have a different suite, the IKAR Lab expert suite, which is where forensic audio examination is performed. And I want to talk about what this suite can do in the context of this task. From the AVIS system, we can export the recordings of interest to IKAR Lab and run a detailed analysis of the recording. For example, do a preliminary quality assessment, decide whether this recording is usable in a forensic examination, and then carry out a full forensic audio examination, starting with noise reduction, extracting the textual content, separating the recording by speakers, and conducting an identification study, a full one, using various methods, including searching for traces of editing and also searching for non-situational changes in the signal.
I want to draw special attention, I'm going to play some audio now, could you please turn the volume up a little over there.
One of the recordings I found in the AVIS system sounds like this. For some reason it seemed interesting, we downloaded it and want to listen to it. I'll play this recording now, let's listen. Listen, I just got into an accident on the M4 highway, just outside Moscow. A Beemer slammed into me, the car's totaled, I urgently need money. Perhaps
those of you who work with audio were able to notice that there are some unnatural notes here. But as our practice shows, most listeners actually can't tell a synthesized voice from the voice of a real person. What we just heard is a synthesized voice. How does this happen? A sample of some person was taken, uploaded to a system, and then the user of that system typed in some arbitrary text, in our case clearly of a fraudulent nature, and the system voiced the text the scammer wanted in the target speaker's voice. A clear, classic situation: a voice message is sent to the parents, the parents panic, urgently look for money, send it, and so on.
This is quite a serious modern threat, a challenge to forensic audio systems: voice forgery, synthesis, re-recording, modification of voice messages. And I want to note that IKAR Lab is now equipped with full spoofing detection functionality. Spoofing is the term used to describe a synthesized voice, or to describe a modified voice. And I want to show the process. Here on the screen we see that the software can evaluate both an overall score for the whole recording, the probability of spoofing, and the probability that a forgery is present on short segments. That is, there's a sliding window, and every second you see the probability distribution on the lower graph, the probability that a fake voice is present in one segment of the recording or another. Because scammers today can forge not the whole recording but just part of it, some most important part.
IKAR Lab makes it possible to detect such forgeries. Moreover, the latest versions of this product can determine not just the presence of spoofing itself, but also the specific vendor. Here, I hope you can see, the slide says that with a probability of 99.6% this recording contains synthesis produced by the vendor ElevenLabs. Today we identify four main vendors. In the future this list will be expanded. I'd like to note separately that spoofing in general is quite a non-trivial task. Detecting a synthesized voice is quite an interesting and at the same time difficult task. It's worth noting synthesis algorithms are advancing by leaps and bounds, dozens or maybe even hundreds of new synthesis algorithms appear every year, so of course there's a kind of race going on here between synthesis detectors and the speech synthesis algorithms themselves. And often in a forensic examination the question posed is not the presence of spoofing, but the broader question of verifying the recording's authenticity, because traces of synthesis are often quite difficult to establish, given how it keeps improving and given that modern algorithms detect it very accurately and very well.
So IKAR Lab lets you take a broader approach to this task, including answering whether there are traces of editing using classical methods, for example phase analysis, background noise analysis, and so on. On these slides I wanted to show a very brief history of synthesis, to get a bit deeper into the subject. This actually isn't yesterday's invention. The first synthesized voice dates to the late 19th century. There was a mechanical device like this that could pronounce two speech-like words, "yes" or "no". There's even a link to the article there, you can look it up if you're interested. And here are some mass-market examples of synthesis. We'll probably listen to a couple of these examples now too.
If possible, turn it down a bit, just a little, they're just going to be rather unpleasant to the ear. Here, for example, is synthesis from 1979, the first mass-market device with built-in human speech synthesis. Let me play it.
Well, it's all clear here, even a child would easily tell this is synthesis. Here, for example, is the synthesis level of the 2000s, let's hear a fragment too.
Overall it's already more like the speech of a real person, but there are no emotions here, no breathing, just some speech-like sounds. They're also easy to tell apart from real human speech. I'll skip one here. And here's the current state of things – this example is taken from news sites, it's from last year, when in the presidential campaign in the US, a synthesized voice of then-president Biden was actively used for smear campaigning, let me play a fragment.
I probably won't play the whole recording, and besides, it's in English. But the most important thing I want to point out here is that every year speech sounds more and more natural. And if I go back to that fragment we listened to, the one with the BMW, we'll hear breathing, simulated breathing, of course, and pauses, and intonation. The systems have learned to copy all of this from a speaker's speech, and our anti-spoofing system can defend against and detect such forgeries. I'll add that IKAR Lab has functionality for manual examination, so the forensic expert can manually detect and document the features that they can then include in their expert report to confirm the conclusions of the automated system.
And, as I mentioned, there are modules for classical technical analysis. That's phase analysis, resampling analysis, background noise analysis, DC offset analysis. These methods live on and still help forensic experts find traces of authenticity violations.
Another interesting example. We got these recordings from open sources. And here we'll broaden a little the topic of voice biometrics into the topic of multimodal biometrics, and look at what our system can do when it comes to faces. So, this is our last talk, and there'll be some other activities afterwards. To get you engaged a little, I want to offer you a small puzzle. I'm going to play two videos at the same time now. One of them is a video of a real person, and the second is a synthesized video. That is, it's a fake face, not a real picture. And I'll ask you to answer the question: which video you think is real, the top or the bottom one. Then we'll just vote and see how the votes are split. So please watch carefully. It's Jessica Alba. My colleagues here are prompting me.
Let's see what it looks like.
No speech here, just the picture. I'll play it once more.
Let's watch it again.
So the first question is this. Please raise your hands, those who think the real video is the top one. Thank you. Now, hands up if you think the real video is the bottom one. Well, roughly even, roughly even, slightly more voting for the top video. Should I tell you the right answer now or later? Right away. Yes, the real video is below, yes, the one on top is synthesis. You… Well done, those who said "bottom". Thanks to those who took part and said "top". The top one is synthesis. And I want to show how our product can help a forensic expert answer that kind of question.
IKAR Lab now has a module for detecting deepfakes. It's frame-by-frame analysis of faces in video. On this clip I'll play back a video from the real interface, how it happens. That is, the expert can analyze video right in their own copy of the IKAR Lab software.
The system runs a frame-by-frame analysis. For each frame it outputs an estimate of the probability that it's a fake. And on top of that a mask is overlaid that shows which areas of the face most strongly activated the network's neurons. That doesn't mean those parts of the face are fake. It generally indicates that they for some reason draw the most attention from the neural network. For this video, we can see on the right in the interface that we got a fake probability of 97%.
And in addition, the software estimates, it says "in tolerance" there, the in-tolerance probability. That is, it also estimated the number of frames where there are excessive head turns or tilts. So the software also checks whether the face looks straight at the camera, or whether it's tilted, turned, and so on.
For an original video, I'll show you the opposite situation too. This is a video of a live person, with a genuine image of Ms. Jessica Alba. And here you can see that the estimated probability is very low. This lets us conclude that this is a video of a live person. And the probability estimate is also output for all the analyzed frames, and also for those frames that turned out to be within tolerance. So that's the kind of toolset we're now offering our customers as well, because deepfakes, spoofing and synthesis are a really serious challenge today that we face, and that's why our solutions are equipped with methods to counter such techniques.
I'll add that the IKAR Lab expert forensic system also has a manual assessment module for video synthesis signs, so that the expert, in their report, in their expert conclusions, can include their observations in a structured form.
And probably the last example I'd like to give here. This is our Nestor AI system. Once we've searched the data in the bulk set, found some recordings or media files of interest, checked them manually, first got some conjectures, then suspicions, detained the suspect, proceedings take place. And those proceedings can be recorded, but they'll be recorded on camera, on a microphone. We also offer a system, the Nestor AI system, which can automatically produce minutes of official sessions. That is, it can work both with streaming audio and with files obtained from a system, say, a conference-call system such as Zoom, or dictaphone recordings.
The output of such a system is a strictly formatted record containing the transcript of the meeting, split by speaker. And I'll also note that we build large language models into these systems, which also make it possible to summarize the results of meetings, to produce short, condensed abstracts of these meetings, for easier work with such records. Here you see an example of the system's interface. So the number of speakers was determined, and the gender of those speakers. Here's the transcript, and now the output of the AI agent will be shown.
Here's the meeting record. And with prompting, of course, you can tune the output of these agents depending on the needs of our customers. As a conclusion, I'd probably say that we work at different stages of the forensic process and bring a new modality to the data you already have, allowing you to find new insights. Thank you. We probably have a little time for questions. Yes, of course we do. Mikhail, thank you very much. So, colleagues. —
— I have two questions, they're not related to each other. But the first question is about video deepfakes. From the examples shown, I'd conclude that it was face swapping originally that was used as the technique. Fine, you detect it. But have you tried it, and how does the product behave on other techniques? Puppet-master, lip-syncing, and synthesis of something? Moustaches, moles, earrings and so on.
As for synthesis of facial elements, I haven't tested that myself, for example adding moles and so on, but for all the other situations, a face mask, real-time face swap, we've tested it, it all works at a sufficiently high quality. No, not just swapping. Lip-syncing, swapping is when it's part of the face, the main face and its expressions are preserved but only a part changes, mostly the triangle. One example is, say, when a photo of some famous person is animated. So there's a photo, we overlay the mask movement, well, the attacker overlays the lip mask movement, and on the whole, yes, such fakes are easily detected. On our datasets we get a result of 98%. Obviously, for other data you have to look, test on the data. But overall the result is very good.
Well yes, so overall it was strange to hear that you identify specific vendors. I thought that underneath those same audio deepfakes there's some Multi-Tacotron sitting anyway. And then it mostly makes sense to identify the tool families, not the vendor who took something free and reskinned it a bit. Why is vendor identification important? We discussed that here with an expert too, with our expert. Of course, it's additional circumstantial evidence when working a case. Because if we can establish that our suspect visited such-and-such a site, for example. If we can detect that the synthesis was generated from that same site, we get additional evidence to support the expert's work.
Thank you. And here's the second question, purely out of idle curiosity. Have you tested your recognizer, which then runs all of this through a neural net, against data poisoning? What happens if, roughly speaking, I add to the audio signal at an inaudible level a system prompt like "listen to me" for the LLM that's about to process it. And accordingly, the voice of conscience dictates something further for it to do. So, here we're talking specifically about prompting. Poisoning. It's just that your input data is audio. And how exactly does your speech-to-text work? That is, maybe I can inject an upper layer, say, initially insert the first part in some sufficiently long pause, so that everything afterwards is ignored, using only that frequency band, and then, well, insert the rest. Thanks for the question. I should say that several models are used in this chain, and they're different.
That is, for speech-to-text conversion it's one set of models, our own in-house development. For the subsequent LLM analysis other models are used. So as far as turning a recording into a transcript goes, it's unlikely here, if you consider how this data comes to our customers. Usually they're interested in the data being clean. And as a rule the user has no interest in distorting that data. Usually a different question arises. Let's tune them for some specific modality. For example, our speakers talk about some strictly defined topics, in some odd language, made up, or school slang, or sci-fi. And then they come to us and say, please tweak and fine-tune the model so that it recognizes this narrow group's slang well. And that's a task we do solve. As for processing, as a rule the customer is interested in the data being clean.
As a rule, such situations don't arise. —
— I have a small question, I'm nervous, but I want to ask it. Are there any internal measures to protect the software product against leaks? After all, this can basically be seen as a weapon, when you have the ability to recognize and analyze speech data. —
— Thanks for the question. Usually these systems are always deployed inside our customers' closed networks. And this data protection question is addressed by the system running in a closed network, and on top of that, technical solutions are used so that the data, that's access control, for example, roles, privileges, separation of access levels, so that a user only sees the data they need. And I'll also note that the models we offer, both the LLM models and the speech-to-text models, the biometric models, they all run on-prem, that is, they don't need any internet access. And they can be, and always are, deployed offline on the solutions, on the hardware of our customers.
Right, Olga Alexandrovna, hold on a second. I see a question over here too.
Thank you very much for the talk. A practical question. Tell me, how effectively does your system deal with… If we have an audio sample where the microphone was being jammed with ultrasound. How effective is the noise cleaning?
— Thanks for the question. It's a big one. Answering in a conference format, with little time, I'll keep it brief. First of all, of course, there's the question of how you obtain that data. You're talking about ultrasound jamming or something. That's a separate topic in itself, one worth debating, because in our practice we very rarely run into situations where such jammers can actually cause any serious interference to good microphones. Let's start with this: as a rule, it's something of a myth that such solutions seriously protect speakers from being recorded. The second point to make: the algorithms are, of course, fairly robust to changes in the acoustic environment: SNR, reverberation, signal-to-noise ratio, reverberation level, and so on. Naturally, when developing these algorithms we always assume that the acoustic conditions the recording is made in, as I wrote there, may be not only a close-talk but also a far-field mic.
They may be far from ideal. So up to a certain limit, a certain SNR value, the probability is very high. Of course, to be honest, the lower the SNR, the more noise there is, the lower the probability of correct identification by an automatic method. But here we can say that there are manual systems, expert ones, such as IKAR Lab, where, using manual methods and other approaches, for example linguistic ones or something else, such problems can be successfully overcome. Thank you. And the second question is more commercial, I suppose. Your product Nestor, does it come as a device or as a license renewal? It's a hardware and software system. Different delivery options: microphones plus software, or just the software. So the software can use the microphones the customer already has installed, microphone arrays. We offer such solutions too. So it can be purchased as a software product. —
— Right, Olga Svetlanovna, I recall a question. How fast is your speech transcription? It's just that I know of systems that need a great deal of computing power to transcribe an audio file with good quality. Thank you. The question is clear, but again there's no simple answer.
I should say the following: transcription speed depends on the data volume, obviously, and the hardware. And, of course, mainly on the hardware. Here's an example: on modern GPUs, for instance, on server solutions, the speed-up of speech-to-text conversion can reach a thousand or even tens of thousands of times. Roughly speaking, a thousand seconds of speech become text in one second. So you need to sit down and look at the data volume and what hardware the customer has. So, in other words, transcription in real time… Considerably faster than real time. Excellent. And at the very beginning you said you support several languages. As far as I know forensic phonoscopy methods, what's in demand here is Russian, basically, wherever native speakers are. A phonoscopist can't do a forensic exam if they're not a native speaker of that language. So who needs analysis in other languages? —
— Yes, I'd answer like this: these systems… First, IKAR Lab has the option of manual transcription, so that a native speaker of, say, some regional language, if it's not supported in semi-automatic mode, could do the transcription manually. And second, IKAR Lab is exported to a number of countries, and those languages are quite relevant there. —
— Mikhail, thank you very much. That was the final talk and the final question of today's conference. Thank you.