> For the complete documentation index, see [llms.txt](https://gitbook.bergelsonlab.com/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://gitbook.bergelsonlab.com/data-pipeline/audio-processing.md).

# Audio Processing

Processing directions from SEEDLingS Wiki--last updated 2016

## Exports

You'll need to create a .csv, a .cha, and an .its file.

* Open **LENA Pro** (elephant insignia) on Lily
* Select the **Client Manager** icon in the upper left-hand corner
* Double-click on the subject that you will be processing
* In the **Report List** window, check if the visit you will be creating exports for, spans one day or two (i.e. if the recorder is on past midnight, it is two days

![Shows only one date per month](https://3364608434-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-LD2B3y86yJyNeihFLqD%2F-LJ_OYolvtgAACKT-KAa%2F-LJ_P-OQCeXTosbmT-vj%2Flena_pic01.1200x1200.png?alt=media\&token=14df8843-e554-445d-823d-9fa49b4e7313)

![Shows two back to back dates for month 7 (adds up to 16h)](https://3364608434-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-LD2B3y86yJyNeihFLqD%2F-LJ_OYolvtgAACKT-KAa%2F-LJ_PAwumENMzxf1go5Y%2Flena_pic02.1200x1200.png?alt=media\&token=0a85b801-39d3-4926-bb18-05146e566d19)

* Select the date(s) that correspond to the recording you want to export and click on the Excel icon in the bottom left-hand corner

![](https://3364608434-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-LD2B3y86yJyNeihFLqD%2F-LJ_OYolvtgAACKT-KAa%2F-LJ_PTz9lLS1UAMdGljr%2Flena_pic03.1200x1200.jpeg?alt=media\&token=2c32a08a-12b6-4d6d-bba0-b3ab128c9be0)

* Make the following selections in the **Export Data** window:
  * Make sure that **5 Minute Detail** is selected under the **Report Elements** heading
  * Under **Specify Dates**, select one or two days

![](https://3364608434-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-LD2B3y86yJyNeihFLqD%2F-LJ_OYolvtgAACKT-KAa%2F-LJ_PrF6if08XXx-oeYC%2Flena_pic04.1200x1200.jpeg?alt=media\&token=edbcd57a-0ab0-4d73-a2de-56ae98c5b09d)

* Under **Export Now**, select CSV
  * Rename the newly exported file as *XX\_XX\_lena5min.csv* (ex: *01\_08\_lena5min.csv*)
  * The lena5min file is a data sheet that chunks the 16 hour audio file into 5 minute segments with speaker categories (ex: child/adult vocalizations, TV/radio/media, distant speech, etc.)
* Select CHA
  * Rename the newly exported file as XX\_XX.cha (ex: 01\_08.cha)
  * This export will create two files (a *.cha* file and a *.wav* file); keep the *.wav* and delete the *.cha*, as we will create one in CLAN later
* Select ITS
  * Rename the newly exported file as XX\_XX.its (ex: 01\_08.its)

## CLAN: .its to .cha

* Now that we have all of our exports in place, we need to convert the *.its* file to a *.cha* file
* Open CLAN and Select **Commands** and change the working directory to where your .its file is saved.
* Write the following command in the window:
  * lena2chat XX\_XX.its

![](https://3364608434-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-LD2B3y86yJyNeihFLqD%2F-LJ_OYolvtgAACKT-KAa%2F-LJ_RR_ar7bq0T6jDZq9%2FScreen%20Shot%202018-08-10%20at%204.05.22%20PM.png?alt=media\&token=a4c807bb-9fe7-4607-814b-7a1a119297c0)

* This makes a "lena.cha" file

## Sound Finder

* There are long periods of time in these recordings where nothing is going on because the child is asleep
* We have a silence finder script that puts this information into the clan file to save coders time (so they're not listening to naps)
* The first step involves finding these long silences in a program called **Audacity**
* Then you edit this list of silent times with a python script looking for little interruptions to the silences so that you can ignore those too (pops or single cries, etc)

![](https://3364608434-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-LD2B3y86yJyNeihFLqD%2F-LJ_OYolvtgAACKT-KAa%2F-LJ_RpTm_mp9bHQo7J-U%2Flena_pic07.1200x1200.jpeg?alt=media\&token=e2f54622-6517-431b-b4de-de3296a03a77)

* Make sure you have pulled the most recent version of **audiowords.py** from the github repository [audiowords](https://github.com/SeedlingsBabylab/audiowords) \[this script requires python (2.x) to function properly]
* Open **Audacity** and go to the same directory you have been using for the above steps, and drag the *.wav* file into your new audacity window
* In the menu bar, select **Analyze** then **Sound Finder**\
  ![](https://mfr.osf.io/export?url=https://osf.io/8na47/?action=download\&direct\&mode=render\&initialWidth=528\&childId=mfrIframe\&format=1200x1200.jpeg)
* In the **Sound Finder** window, change your settings to match those below

![Change Minimum duration of silence last, as changes in the others will affect it](https://3364608434-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F-LD2B3y86yJyNeihFLqD%2Fuploads%2FDHaTRqQrxWtALJYFE7qZ%2Fsound-finder_settings.png?alt=media\&token=253b5ccd-4421-41c2-befd-9d14f90f6319)

### Guidelines for Processing

1. Here are some general guidelines to follow when scanning through the audio file within **Audacity**
2. There likely be 2 to 3 regions of minimal activity (naps, sleeping, possibly a car ride, etc)

   a. Pay close attention to these regions as they may contain some verbal production that the script did not pick up based on its parameters

   b. **Caveat:** Only keep the bits of these regions where there is verbal production (concrete object words); if the speaker is soothing the baby back to sleep that does not have to be included as it would not be coded later on
3. Scan through the stretches of sound, focusing on moments with lower peaks of activity or bits with constant amplitudes (these areas denote a possible car ride or noise/sound machine, for example, that will help you indicate the onsets/offsets of naps, television, radio, etc.)
4. Take care when zooming into different regions (*Ctrl+Scroll*) where you think think there is little to no activity, as sometimes there may be!

**Remember:** This is a preliminary pass whose main purpose is to save time for the in-depth coding that will be performed in CLAN. Keep this in mind when processing, as it will help you when making decisions on what bits to keep

### Exporting

1. Export the sound segments
2. In the menu bar, select **File** then **Export Labels**
3. Rename the new file as *Label\_Track.txt* and save it to the same directory

### Audiowords: Running the Script

1. We now have all the files that we need to complete processing
2. In order to do so, we will use a script called *audiowords.py* from [this](https://github.com/SeedlingsBabylab/audiowords) repo.

Click **Load All (cha)** and select the main CLAN file (ex, 01\_08.lena.cha)&#x20;

**Info:** The program will load and generate the other files that are necessary, running through all the steps at once. It assumes that all the necessary files are within the same directory as the original CLAN file that was loaded. It will output the *XX\_XX\_silences.txt* regions, *XX\_XX\_silences\_added.cha*, and *XX\_XX\_subregions.cha* exports to this same directory.

![](https://3364608434-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-LD2B3y86yJyNeihFLqD%2F-LJ_OYolvtgAACKT-KAa%2F-LJ_V7rZO5ZFFolWAbv-%2Flena_pic11.1200x1200.jpeg?alt=media\&token=de8613e0-3b2f-4d4b-9ce5-2678ebb6599a)

The format it's expecting files to be in:

* 01\_08.lena.cha
* 01\_08\_lena5min.csv
* Label\_Track.txt
* 01\_08\_silences.txt **(output)**
* 01\_08\_silences\_added.cha **(output)**
* 01\_08\_subregions.cha **(output)**

After the script is finished running, the terminal window will give you the following information:

* The number of silences which should match up with the *XX\_XX\_silences.txt* file that was just generated
* Whether or not the file spans over one day or two
* An issue within the CLAN (*XX\_XX\_subregions.cha*) file where the onset is greater than the offset; **this is ultra important and will be surrounded by asterisks in the terminal window**
  * Open the *XX\_XX\_subregions.cha* file and use the **Esc+L** function to search for the **Line Number** specified in the terminal window
  * To make the correction; open timestamps (**Esc+a**) and change the *bad offset* to match the onset of the next line
  * Insert the following as comment under the line you just changed: **%com: OR %xcom: (tab) manually adjusted timestamp**
  * Run the **Esc+L** command to check any remaining issues in the file that may have flown under the radar
