In brief
The closing talk of the conference, a vendor presentation by genre: three STC products along the chain "raw audio → search → forensic examination → minutes" — AVIS, IKAR Lab, Nestor AI. All the figures come from the demos. The gist: a synthesized voice is indistinguishable by ear, dozens and hundreds of new algorithms appear every year, and therefore a synthesis detector is unreliable on its own — the question posed is the authenticity of the recording, with the classical methods and manual documenting of the features. The questions from the audience turned out more substantive than the talk.
Key points
- STC: over 35 years, more than 5,000 projects, the NIST and CHiME Challenge competitions; on LLMs — GigaChat and freely available models.
- The biometrics are language-independent and text-independent, and there are manual methods too: the conclusions of the automated system need expert confirmation. More than 70 parts of the body are involved in the voice, phonoscopy has been around for more than 30 years — a forensic examination, identification by the formants and the fundamental frequency.
- The input — interception, microphone recordings, voice messages from a seized device; AVIS — analysis, "the IKAR Lab 3 forensic suites" — evidence (whether the recording is usable, noise reduction, speaker separation, identification, editing), Nestor AI — the minutes.
- AVIS: gender and language, recognition in 18 languages (Russian, Ukrainian, Kazakh, Arabic), biometric search, translation, graphs; LLMs give a summary and find "code words and slang".
- The demo — a fraudulent voice message about an accident on the M4; spoofing is scored both for the whole recording and second by second with a sliding window: they forge only a part.
- The synthesis vendor: 99.6% and ElevenLabs in the demo, four in all; "dozens or maybe even hundreds" of new algorithms a year — hence the question of authenticity and the classical methods: phase, resampling, background noise, DC offset.
- The history of synthesis: the "yes"/"no" of the late 19th century, a 1979 device, 2000s synthesis with no emotions, last year's synthesized voice of Biden; today's copies breathing and intonation.
- The Jessica Alba video: the room split "roughly even", the top one was the synthesis. The deepfake module: a frame-by-frame probability, a mask of what "activated the network's neurons" (the network's attention, not the forgery), an "in tolerance" figure by head tilts; 97% on the fake.
- Nestor AI: streaming audio and files (Zoom, a dictaphone), strictly formatted minutes split by speaker, LLM abstracts, an AI agent.
Tools, artifacts, technologies
- Offered to customers: AVIS (STC) — search and analytics; IKAR Lab ("IKAR Lab 3" once) — forensic examination, spoofing, video deepfakes, manual modules, export to a number of countries; Nestor AI — a hardware and software system for minutes.
- In use: GigaChat, freely available LLMs, their own speech-to-text; NIST and the CHiME Challenge — competitions; ElevenLabs — the vendor from the demo (the other three are not named); Multi-Tacotron — named from the audience.
- Zoom and dictaphones — the input for Nestor AI; server GPUs — the speed-up of transcription. Deepfake techniques (from the audience): face swapping, puppet-master, lip-syncing, synthesis of moustaches, moles, earrings, a photo being "animated"; ultrasonic jammers — criticized.
Legal and organizational context
No laws, articles of law or agencies were named; the terminology is procedural — "forensic speaker examination", "expert report", "lawful interception". Phonoscopy has for more than 30 years been a type of forensic examination in Russia and abroad; the conclusions of the automated system require expert confirmation, hence the modules for manually documenting the features for the report. AVIS — search, IKAR Lab — evidence. Deployment — the customer's closed network, on-prem, offline, roles and access levels.
Questions from the audience
The audience had no microphone and there is no speaker labeling; no names were given, attribution is from the context.
- Other deepfake techniques — puppet-master, lip-sync, moles? (the same man — the next two as well) → He has not tested facial elements; a photo being "animated" — 98% on STC's own datasets.
- Why the vendor, if underneath it is "some Multi-Tacotron" anyway? → Circumstantial evidence: the suspect's site and where the generation happened.
- Prompt injection through audio? → The models in the chain are different; unlikely, "the customer is interested… in the data being clean".
- Protecting the product against leaks? → A closed network, on-prem.
- Noise cleaning when the microphone is jammed with ultrasound? → "Something of a myth"; beyond the SNR limit — manual methods.
- Is Nestor a device or a license? → A hardware and software system: microphones with software, or software alone.
- Transcription speed? ("Olga") → On server GPUs, 1,000 seconds of speech in one second.
- Forensic examinations in other languages? (the same woman) → Manual transcription by a native speaker; export to a number of countries.
The speaker's position
The tone of a vendor presentation: "internationally recognised", "it all works at a sufficiently high quality". At the same time the automation is not presented as a replacement for the expert: its conclusions require confirmation, the classical authenticity methods work, synthesis detectors are losing the race to the algorithms. He argues with the audience about the synthesis vendor: the opponent — what has to be identified is the tool families, the speaker — the vendor as evidence. Some of the limitations he acknowledges himself: he has not tested facial elements, the 98% is on his own datasets, as the SNR drops the accuracy drops.
Quotes
- "…the conclusions of any automatic system need expert confirmation."
- "…most listeners actually can't tell a synthesized voice from the voice of a real person."
- "…it's something of a myth that such solutions seriously protect speakers from being recorded."
- "…the more noise there is, the lower the probability of correct identification…"