PofoliaShared via Pofolia

· 2021

AST: Audio Spectrogram Transformer

Yuan Gong, Yu-An Chung, James Glass

Short summary

The Audio Spectrogram Transformer (AST) is the first purely attention-based model for audio classification, achieving new state-of-the-art results: 0.485 mAP on AudioSet, 95.6% accuracy on ESC-50, and 98.1% accuracy on Speech Commands V2.

AI-generated from the title and abstract; the full text is not read.

TakeawaysIn the app
Key pointsIn the app
Ask the paperIn the app

The rest is in the Pofolia app

Takeaways, key points and questions to the paper; new summaries every day for your field. Free.

Sign in on the web to open

Field: Signal Processing

Signal ProcessingComputer Science