[{"data":1,"prerenderedAt":705},["ShallowReactive",2],{"$f2kp100t9t3lju":3},{"title":4,"date":5,"tags":6,"categories":9,"draft":11,"_id":12,"slug":13,"path":14,"document":15,"excerpt":703,"readingTimeMinutes":704},"Easily Transcribe Podcasts with Whisper.cpp","2024-01-08T00:00:00.000Z",[7,8],"whisper.cpp","ml",[10],"programming",false,"content:blog:podcast-transcription-whispercpp","podcast-transcription-whispercpp","\u002Fblog\u002Fpodcast-transcription-whispercpp",{"frontmatter":16,"meta":17,"nodes":18},{},{},[19,38,42,47,69,134,140,227,234,240,312,315,410,413],[20,21,22,23,27,28,32,33,37],"p",{},"If you've ever had the need to transcribe a podcast, lecture, or some other audio recording, it turns out it's surprisingly easy with the extremely impressive ",[24,25,7],"a",{"href":26},"https:\u002F\u002Fgithub.com\u002Fggerganov\u002Fwhisper.cpp"," project. This high-performance fork of ",[24,29,31],{"href":30},"https:\u002F\u002Fgithub.com\u002Fopenai\u002Fwhisper","OpenAI's Whisper"," can run on all sorts of hardware -- including my M1 Mac Mini. Let's walk through an example from start-to-finish of transcribing an episode of the ",[24,34,36],{"href":35},"https:\u002F\u002Fpodcasts.apple.com\u002Fus\u002Fpodcast\u002Falter-everything\u002Fid1356137854","Alter Everything"," podcast.",[39,40,41],null,{},"more",[43,44,46],"h2",{"id":45},"obtain-audio-files","Obtain Audio File(s)",[20,48,49,50,54,55,58,59,61,62,64,65,68],{},"First, let's get the ",[51,52,53],"code",{},"wav"," file from YouTube using the ",[51,56,57],{},"youtube-dl"," utility. It should be noted that ",[51,60,7],{}," expects ",[51,63,53],{}," filetypes, and this utility defaults to ",[51,66,67],{},"mp3",".",[70,71,74],"pre",{"language":72,"class":73},"sh","shiki shiki-themes tokyo-night dark:tokyo-night",[51,75,77,94,95,94,103,94,113,94,123],{"class":76},"language-sh",[78,79,82,86,90],"span",{"class":80,"style":81},"line","display: inline",[78,83,85],{"style":84},"color:#C0CAF5"," $",[78,87,89],{"style":88},"color:#9ECE6A"," youtube-dl",[78,91,93],{"style":92},"color:#89DDFF"," \\","\n",[78,96,97,101],{"class":80,"style":81},[78,98,100],{"style":99},"color:#E0AF68","    --extract-audio",[78,102,93],{"style":92},[78,104,105,108,111],{"class":80,"style":81},[78,106,107],{"style":99},"    --audio-format",[78,109,110],{"style":88}," wav",[78,112,93],{"style":92},[78,114,115,118,121],{"class":80,"style":81},[78,116,117],{"style":99},"    --output",[78,119,120],{"style":88}," podcast.wav",[78,122,93],{"style":92},[78,124,125,128,131],{"class":80,"style":81},[78,126,127],{"style":92},"    \"",[78,129,130],{"style":88},"https:\u002F\u002Fwww.youtube.com\u002Fwatch?v=CoUN690wSYQ",[78,132,133],{"style":92},"\"",[20,135,136,137,139],{},"This file has a 44.1 kHz sample rate, and ",[51,138,7],{}," expects 16 kHz, so let's go ahead and convert that.",[70,141,142],{"language":72,"class":73},[51,143,144,94,153,94,165,94,167,94,189,94,191,94,199,94,209,94,211,94,217,94,222],{"class":76},[78,145,146,148,151],{"class":80,"style":81},[78,147,85],{"style":84},[78,149,150],{"style":88}," file",[78,152,120],{"style":88},[78,154,155,158,161],{"class":80,"style":81},[78,156,157],{"style":84},"podcast.wav:",[78,159,160],{"style":88}," RIFF",[78,162,164],{"style":163},"color:#A9B1D6"," (little-endian) data, WAVE audio, Microsoft PCM, 16 bit, stereo 44100 Hz",[78,166],{"class":80,"style":81},[78,168,169,171,174,177,179,182,186],{"class":80,"style":81},[78,170,85],{"style":84},[78,172,173],{"style":88}," ffmpeg",[78,175,176],{"style":99}," -i",[78,178,120],{"style":88},[78,180,181],{"style":99}," -ar",[78,183,185],{"style":184},"color:#FF9E64"," 16000",[78,187,188],{"style":88}," podcast-16khz.wav",[78,190],{"class":80,"style":81},[78,192,193,195,197],{"class":80,"style":81},[78,194,85],{"style":84},[78,196,150],{"style":88},[78,198,188],{"style":88},[78,200,201,204,206],{"class":80,"style":81},[78,202,203],{"style":84},"podcast-16khz.wav:",[78,205,160],{"style":88},[78,207,208],{"style":163}," (little-endian) data, WAVE audio, Microsoft PCM, 16 bit, stereo 16000 Hz",[78,210],{"class":80,"style":81},[78,212,213],{"class":80,"style":81},[78,214,216],{"style":215},"color:#51597D;--shiki-light-font-style:italic","# NOTE: it looks like it's possible to specify this conversion as a post-process as a",[78,218,219],{"class":80,"style":81},[78,220,221],{"style":215},"# flag to the `youtube-dl` command -- I will explore this further next time...",[78,223,224],{"class":80,"style":81},[78,225,226],{"style":215},"# youtube-dl --extract-audio --audio-quality 0 --audio-format mp3 --postprocessor-args \"-ar 44100\" %dl%",[43,228,230,231,233],{"id":229},"build-whispercpp-transcribe-audio","Build ",[51,232,7],{}," & Transcribe Audio",[20,235,236,237,239],{},"Then, let's get the latest version of ",[51,238,7],{},", download the English Whisper model, and build the example.",[70,241,242],{"language":72,"class":73},[51,243,244,94,249,94,278,94,280,94,285,94,298,94,300,94,305],{"class":76},[78,245,246],{"class":80,"style":81},[78,247,248],{"style":215},"# Clone the `whisper.cpp` repository",[78,250,251,253,256,259,262,265,268,271,275],{"class":80,"style":81},[78,252,85],{"style":84},[78,254,255],{"style":88}," git",[78,257,258],{"style":88}," clone",[78,260,261],{"style":99}," --depth",[78,263,264],{"style":184}," 1",[78,266,267],{"style":88}," git@github.com:ggerganov\u002Fwhisper.cpp",[78,269,270],{"style":92}," &&",[78,272,274],{"style":273},"color:#0DB9D7"," cd",[78,276,277],{"style":88}," whisper.cpp",[78,279],{"class":80,"style":81},[78,281,282],{"class":80,"style":81},[78,283,284],{"style":215},"# Download the English Whisper model in `ggml` format",[78,286,287,289,292,295],{"class":80,"style":81},[78,288,85],{"style":84},[78,290,291],{"style":88}," bash",[78,293,294],{"style":88}," .\u002Fmodels\u002Fdownload-ggml-model.sh",[78,296,297],{"style":88}," base.en",[78,299],{"class":80,"style":81},[78,301,302],{"class":80,"style":81},[78,303,304],{"style":215},"# Build the main example",[78,306,307,309],{"class":80,"style":81},[78,308,85],{"style":84},[78,310,311],{"style":88}," make",[20,313,314],{},"And finally, let's transcribe that podcast!",[70,316,317],{"language":72,"class":73},[51,318,319,94,328,94,338,94,348,94,355,94,363,94,365,94,370,94,375,94,380,94,385,94,390,94,395,94,400,94,405],{"class":76},[78,320,321,323,326],{"class":80,"style":81},[78,322,85],{"style":84},[78,324,325],{"style":88}," .\u002Fmain",[78,327,93],{"style":92},[78,329,330,333,336],{"class":80,"style":81},[78,331,332],{"style":99},"    -m",[78,334,335],{"style":88}," ~\u002Fworkspace\u002Fwhisper.cpp\u002Fmodels\u002Fggml-base.en.bin",[78,337,93],{"style":92},[78,339,340,343,346],{"class":80,"style":81},[78,341,342],{"style":99},"    -f",[78,344,345],{"style":88}," ~\u002FDownloads\u002Fpodcast-16khz.wav",[78,347,93],{"style":92},[78,349,350,353],{"class":80,"style":81},[78,351,352],{"style":99},"    --output-vtt",[78,354,93],{"style":92},[78,356,357,360],{"class":80,"style":81},[78,358,359],{"style":99},"    --output-file",[78,361,362],{"style":88}," out",[78,364],{"class":80,"style":81},[78,366,367],{"class":80,"style":81},[78,368,369],{"style":215},"# whisper_print_timings:     load time =   114.71 ms",[78,371,372],{"class":80,"style":81},[78,373,374],{"style":215},"# whisper_print_timings:     fallbacks =   0 p \u002F   0 h",[78,376,377],{"class":80,"style":81},[78,378,379],{"style":215},"# whisper_print_timings:      mel time =   692.20 ms",[78,381,382],{"class":80,"style":81},[78,383,384],{"style":215},"# whisper_print_timings:   sample time = 22278.10 ms \u002F 27893 runs (    0.80 ms per run)",[78,386,387],{"class":80,"style":81},[78,388,389],{"style":215},"# whisper_print_timings:   encode time = 10000.75 ms \u002F    55 runs (  181.83 ms per run)",[78,391,392],{"class":80,"style":81},[78,393,394],{"style":215},"# whisper_print_timings:   decode time =   331.77 ms \u002F    54 runs (    6.14 ms per run)",[78,396,397],{"class":80,"style":81},[78,398,399],{"style":215},"# whisper_print_timings:   batchd time = 45236.73 ms \u002F 27566 runs (    1.64 ms per run)",[78,401,402],{"class":80,"style":81},[78,403,404],{"style":215},"# whisper_print_timings:   prompt time =  1921.90 ms \u002F 11832 runs (    0.16 ms per run)",[78,406,407],{"class":80,"style":81},[78,408,409],{"style":215},"# whisper_print_timings:    total time = 80709.54 ms",[20,411,412],{},"A full podcast transcribed in ~80 seconds on an M1 Mac Mini -- not too bad!",[70,414,415],{"class":73},[51,416,417,94,422,94,427,94,432,94,437,94,441,94,446,94,451,94,455,94,460,94,465,94,469,94,474,94,479,94,483,94,488,94,493,94,497,94,502,94,507,94,511,94,516,94,521,94,525,94,530,94,535,94,539,94,544,94,549,94,553,94,558,94,563,94,567,94,572,94,577,94,581,94,586,94,591,94,595,94,600,94,605,94,609,94,614,94,619,94,623,94,628,94,633,94,637,94,642,94,647,94,651,94,656,94,661,94,665,94,670,94,675,94,679,94,684,94,689,94,693,94,698],{},[78,418,419],{"class":80,"style":81},[78,420,421],{},"# out.vtt",[78,423,424],{"class":80,"style":81},[78,425,426],{},"",[78,428,429],{"class":80,"style":81},[78,430,431],{},"00:00:00.000 --> 00:00:06.480",[78,433,434],{"class":80,"style":81},[78,435,436],{}," >> Hi everyone. We recently launched a short engagement feedback survey for the Alter Everything",[78,438,439],{"class":80,"style":81},[78,440,426],{},[78,442,443],{"class":80,"style":81},[78,444,445],{},"00:00:06.480 --> 00:00:11.360",[78,447,448],{"class":80,"style":81},[78,449,450],{}," podcast. Click the link in the episode description wherever you're listening to let us know what",[78,452,453],{"class":80,"style":81},[78,454,426],{},[78,456,457],{"class":80,"style":81},[78,458,459],{},"00:00:11.360 --> 00:00:16.320",[78,461,462],{"class":80,"style":81},[78,463,464],{}," you think and help us improve our show.",[78,466,467],{"class":80,"style":81},[78,468,426],{},[78,470,471],{"class":80,"style":81},[78,472,473],{},"00:00:16.320 --> 00:00:21.200",[78,475,476],{"class":80,"style":81},[78,477,478],{}," Welcome to Alter Everything, a podcast about data science and analytics culture. I'm Megan",[78,480,481],{"class":80,"style":81},[78,482,426],{},[78,484,485],{"class":80,"style":81},[78,486,487],{},"00:00:21.200 --> 00:00:26.440",[78,489,490],{"class":80,"style":81},[78,491,492],{}," Dibble and today I'm talking with Nick Schrock, CTO and founder of Dagster Labs. We discussed",[78,494,495],{"class":80,"style":81},[78,496,426],{},[78,498,499],{"class":80,"style":81},[78,500,501],{},"00:00:26.440 --> 00:00:31.560",[78,503,504],{"class":80,"style":81},[78,505,506],{}," data engineering trends, challenges in the field, why he started his company, and what",[78,508,509],{"class":80,"style":81},[78,510,426],{},[78,512,513],{"class":80,"style":81},[78,514,515],{},"00:00:31.560 --> 00:00:38.960",[78,517,518],{"class":80,"style":81},[78,519,520],{}," makes him excited about the future of data engineering. Let's get started.",[78,522,523],{"class":80,"style":81},[78,524,426],{},[78,526,527],{"class":80,"style":81},[78,528,529],{},"00:00:38.960 --> 00:00:42.720",[78,531,532],{"class":80,"style":81},[78,533,534],{}," >> Hi, Nick. It's great to have you on our show today. Thanks for being here.",[78,536,537],{"class":80,"style":81},[78,538,426],{},[78,540,541],{"class":80,"style":81},[78,542,543],{},"00:00:42.720 --> 00:00:43.920",[78,545,546],{"class":80,"style":81},[78,547,548],{}," >> Thanks for having me.",[78,550,551],{"class":80,"style":81},[78,552,426],{},[78,554,555],{"class":80,"style":81},[78,556,557],{},"00:00:43.920 --> 00:00:48.280",[78,559,560],{"class":80,"style":81},[78,561,562],{}," >> Yeah. Could you start off by giving an introduction to yourself for our listeners?",[78,564,565],{"class":80,"style":81},[78,566,426],{},[78,568,569],{"class":80,"style":81},[78,570,571],{},"00:00:48.280 --> 00:00:52.920",[78,573,574],{"class":80,"style":81},[78,575,576],{}," >> Sure. My name is Nick Schrock. I'm the CTO and founder of Dagster Labs. There's the",[78,578,579],{"class":80,"style":81},[78,580,426],{},[78,582,583],{"class":80,"style":81},[78,584,585],{},"00:00:52.920 --> 00:00:59.520",[78,587,588],{"class":80,"style":81},[78,589,590],{}," company behind Dagster, which is a data orchestration framework. Prior to doing this, I was an engineer",[78,592,593],{"class":80,"style":81},[78,594,426],{},[78,596,597],{"class":80,"style":81},[78,598,599],{},"00:00:59.520 --> 00:01:05.960",[78,601,602],{"class":80,"style":81},[78,603,604],{}," at Facebook from 2009, 2017. While I was there, I found a team called product infrastructure",[78,606,607],{"class":80,"style":81},[78,608,426],{},[78,610,611],{"class":80,"style":81},[78,612,613],{},"00:01:05.960 --> 00:01:09.800",[78,615,616],{"class":80,"style":81},[78,617,618],{}," whose goal was to make our application developers more efficient and productive, and a bunch",[78,620,621],{"class":80,"style":81},[78,622,426],{},[78,624,625],{"class":80,"style":81},[78,626,627],{},"00:01:09.800 --> 00:01:13.840",[78,629,630],{"class":80,"style":81},[78,631,632],{}," of open source work came out of that actually, one of which was React, which I had nothing",[78,634,635],{"class":80,"style":81},[78,636,426],{},[78,638,639],{"class":80,"style":81},[78,640,641],{},"00:01:13.840 --> 00:01:18.040",[78,643,644],{"class":80,"style":81},[78,645,646],{}," to do with, but actually the CEO of Dagster Labs co-created and I personally co-created",[78,648,649],{"class":80,"style":81},[78,650,426],{},[78,652,653],{"class":80,"style":81},[78,654,655],{},"00:01:18.040 --> 00:01:22.640",[78,657,658],{"class":80,"style":81},[78,659,660],{}," GraphQL. So as I like to say, Pete and I were present at the creation of the full hipster",[78,662,663],{"class":80,"style":81},[78,664,426],{},[78,666,667],{"class":80,"style":81},[78,668,669],{},"00:01:22.640 --> 00:01:28.680",[78,671,672],{"class":80,"style":81},[78,673,674],{}," stack. I moved on to Facebook in 2017, figuring out what to do next, and this data engineering",[78,676,677],{"class":80,"style":81},[78,678,426],{},[78,680,681],{"class":80,"style":81},[78,682,683],{},"00:01:28.680 --> 00:01:32.960",[78,685,686],{"class":80,"style":81},[78,687,688],{}," and data orchestration problem really got me hooked actually quite soon after I left,",[78,690,691],{"class":80,"style":81},[78,692,426],{},[78,694,695],{"class":80,"style":81},[78,696,697],{},"00:01:32.960 --> 00:01:36.280",[78,699,700],{"class":80,"style":81},[78,701,702],{}," and the rest is history. I'm sure we'll get into that more.","If you've ever had the need to transcribe a podcast, lecture, or some other audio recording, it turns out it's surprisingly easy with the extremely impressive whisper.cpp project. This highperformance fork of OpenAI's Whisper can run on all sorts of hardware including my M1 Mac Mini. Let's walk through an example from starttofinish of transcribing an episode of the Alter Everything podcast.",1,1790648870529]