In brief

The closing talk of the conference, a vendor presentation by genre: three STC products along the chain "raw audio → search → forensic examination → minutes" — AVIS, IKAR Lab, Nestor AI. All the figures come from the demos. The gist: a synthesized voice is indistinguishable by ear, dozens and hundreds of new algorithms appear every year, and therefore a synthesis detector is unreliable on its own — the question posed is the authenticity of the recording, with the classical methods and manual documenting of the features. The questions from the audience turned out more substantive than the talk.

Key points

Tools, artifacts, technologies

No laws, articles of law or agencies were named; the terminology is procedural — "forensic speaker examination", "expert report", "lawful interception". Phonoscopy has for more than 30 years been a type of forensic examination in Russia and abroad; the conclusions of the automated system require expert confirmation, hence the modules for manually documenting the features for the report. AVIS — search, IKAR Lab — evidence. Deployment — the customer's closed network, on-prem, offline, roles and access levels.

Questions from the audience

The audience had no microphone and there is no speaker labeling; no names were given, attribution is from the context.

The speaker's position

The tone of a vendor presentation: "internationally recognised", "it all works at a sufficiently high quality". At the same time the automation is not presented as a replacement for the expert: its conclusions require confirmation, the classical authenticity methods work, synthesis detectors are losing the race to the algorithms. He argues with the audience about the synthesis vendor: the opponent — what has to be identified is the tool families, the speaker — the vendor as evidence. Some of the limitations he acknowledges himself: he has not tested facial elements, the 98% is on his own datasets, as the SNR drops the accuracy drops.

Quotes