Showing posts with label techniques. Show all posts
Showing posts with label techniques. Show all posts

Nov 17, 2008

Basic Encoding Techniques

Whether you encode your podcast by exporting directly from your editing platform or by using a stand-alone encoder, you can specify a number of parameters. You may have only a few choices if you're using encoding presets, or you may have the opportunity to specify exactly how you want your podcast encoded.

In the early days of low bit-rate encoding, back when people were connected to the Internet via slow modems, encoding technology was limited and required lots of tweaking to extract the best quality. Now, ten years later, codec technology and Internet connection speeds have improved so much that encoding high-quality podcasts should be within everyone's reach.

This is particularly true of audio podcasts. Modern codecs such as RealAudio and Windows Media Audio are capable of attaining FM-mono quality at a mere 32 kbps. The MP3 codec lags behind in quality, but because you can safely encode your podcast at 128 kbps, you should not have any quality issues.

Video is a little trickier. Assuming the majority of your audience is on a broadband connection, your video quality is limited by available bandwidth. Although you can't expect DVD quality at these bit rates, there's no reason why you can't create a perfectly acceptable video experience. This chapter helps you choose settings that should do the job. Let's start off with the easy stuff — audio encoding.

Audio Encoding


Audio encoding is easy, for a number of reasons. Raw audio files are large, but nowhere near as huge as video files. Therefore, the amount of compression that is needed to reduce them to a size that is suitable for Internet distribution is not excessive. Audio codec technology has progressed to a point where low bit rate encoding produces very good results. Podcasting reaps the benefits of ten years of cutthroat competition between RealNetworks and Microsoft, and the progress made by the MPEG organization with AAC encoding.

Because modern codecs sound so good, you really don't need to do much tweaking when you're encoding audio. You really have to decide only three things: whether to encode in stereo or mono, whether to use a speech or a music codec, and what bit rate to use.

Mono versus stereo
The first thing to decide is whether to encode your podcast in stereo or mono. If your program is predominantly interviews or spoken word, encode in mono. Mono encodings are always higher fidelity at a given bit rate, because only a single channel is encoded instead of two. If you're encoding in mono, you can use a lower bit rate and get the same quality or you can get better quality than a stereo encoding at the same bit rate.

If your content is predominantly music, you should encode in stereo, although it isn't strictly necessary. Even though music is recorded in stereo, most of the content is right in the center of the mix. The lead vocal, the snare drum, the bass drum, all will be right in the center of the speakers. And watch where you place your speakers. If you aren't sitting directly between the speakers, you aren't experiencing the full stereo effect anyway. However, one good reason to target stereo if you're playing music is that half your audience may be listening on headphones, which exaggerates the stereo effect.

Speech versus music
The next thing to decide is whether to use a speech codec or a music codec. If you're encoding an MP3 file, you don't have a choice. MP3 is a music codec. The good news is that MP3 is perfectly suitable as a speech codec as well, provided the bit rate is high enough.

Speech codecs can take special shortcuts during the encoding process due to the nature of speech content. With speech, the dynamic range tends to be very limited, as is the frequency range. After you start talking, the chances are good that you'll continue to speak at roughly the same volume and in the same register. Knowing this, a speech codec can make intelligent decisions about how to encode the audio.

Music content, on the other hand, has a wide dynamic and frequency range. There are bass drums and bass guitars, as well as crashing cymbals and violins. The shortcuts that a speech codec takes are completely unsuitable for encoding music content.

So the choice is fairly obvious: If you're encoding content that is speech only, you can encode at very low bit rates and still achieve high quality using a speech codec. However, for most applications, a music codec is perfectly appropriate.

Bit rates, sample rates, and quality equivalents
The most important decision to make about your audio podcast encoding is what bit rate to use. The bit rate determines the eventual file size of your podcast, which in turn determines how long it takes to download. The bit rate also determines the fidelity of your podcast. The higher the bit rate, the higher fidelity your podcast is.

The listed audio bit rates range from 20 kbps to 256 kbps. If you're producing audio-only podcasts, you should target somewhere between 64 kbps and 128 kbps. If you're encoding predominantly speech, you can safely stay at the low end of that; if you're encoding music, you may want to stick to the higher end of the spectrum.

Note At the end of the day, you know best how you want the podcast to sound. Try encoding at a couple of different bit rates, and see which one sounds best to you.

The other thing you may be able to set is the sampling rate. The sampling rate determines how much high-frequency information is encoded. For example, CD-quality audio uses a sample rate of 44.1 KHz, to capture the full 20–20,000 Hz frequency range. The sampling rate has to be at least double the highest frequency you're trying to capture. Depending on what bit rate you're targeting, you may be offered a few different sampling rates.

The interesting thing about sampling rates is that a higher sampling rate isn't necessarily better. The sampling rate determines how often the incoming audio signal is sampled, so it determines how much audio the encoder has to try to encode. If you set a higher sampling rate, you're telling the encoder to try to encode more high-frequency information, but the encoder may have to sacrifice the overall quality of the encoding. Essentially, the sampling rate determines the trade-off between the frequency range and the fidelity of the encoding. At a given bit rate, an encoder can offer higher fidelity with a reduced frequency range or reduced fidelity with a higher frequency range.

We suggest that you choose a lower sampling rate, thereby allowing the encoder to create a higher fidelity version of your podcast. There is very little information above 16 KHz in most audio programming, and most people don't have speakers that reproduce it faithfully anyway. Therefore, choosing a 32 KHz or 22 KHz sampling rate should provide more than enough high-frequency information.

Aug 5, 2008

Advanced Video Production Techniques - Adding Titles

Most professional video programming has some sort of opening sequence that usually includes lots of candid footage mixed with shots of the star(s) and some sort of graphic rendition of the title of the program. You should take the same approach. If your show has a name, let folks know about it! If they download it to their iPod and forget about it until it magically appears on their screen one day when they're browsing through their clips library, you want them to know the name of the program and who you are. So you'll probably want to use titles.

However, the problem is that what looks good (and is legible) on a television screen in general ends up way too small to be read on a small 320×240 screen. Titles at the bottom of the screen (called lower thirds) can be very hard to read if they're not done with large enough fonts. PowerPoint slides are particularly tough, because most people try to pack far too much information into a single slide, which makes it difficult for people to absorb, and the small fonts become very hard to read when reduced. To top it all off, video codecs have a tough time with text, because they don't treat it as being distinct from the video. So when your podcast is encoded, you're going to lose even more quality, as depicted in Figure 10.9.



Figure 1: PowerPoint slides are a good example of why text is tough: (a) Scaled to 320×240 and (b) after encoding at 300 Kbps.


The PowerPoint slide in Figure 1 isn't too bad to start off with; it has only five main points on the slide. By the time the slide is reduced to 320×240, the sub-points are too hard to read, and after the encoding process, even the main points are starting to look a little ragged.

If you're going to use text in your podcast, think big. Try not to have more than three or four points per slide if you're using PowerPoint, and if you're adding titles to your show and/or your guests, make sure to use a font large enough so that it is legible after the encoding process.

Jun 4, 2008

Camera Techniques

After you've taken some time to consider and light your subject, you're ready for the "camera" part of the "lights … camera … action" cliché. As mentioned earlier, your camera is probably the most important part of your video production chain, because if your camera doesn't faithfully render your perfectly lit scene, you're starting off with compromised quality, which propagates quality issues throughout your entire video podcast.

In many ways, shooting a video podcast should be no different than shooting for broadcast. You're trying to get the best shot, with plenty of light and color information and lots of detail. Not only does this look best when you're shooting, but it also makes for a better-looking podcast. However, you should take into account a number of things, because the Internet isn't quite ready for primetime, and podcasts are watched on computer screens and portable media players. Bearing this in mind, you should consider things like shot composition and what camera moves you have planned, because they have a direct affect on the quality of your podcast.

Shot composition
The most obvious thing to think about is shot composition. In most cases, your podcast will end up as a relatively small screen resolution, probably 320×240. An iPod screen measures about 2 inches wide by 1.5 inches tall. On a computer monitor, depending on how your resolution is set, this same resolution can be up to roughly 4 inches wide by 3 inches tall. Either way you look at it, it's not the largest screen in the world. Therefore, you probably want to do away with your long shots and concentrate on medium shots and close-ups.

Because podcasting tends to be a very personal medium, the most common video podcasts tend to concentrate on "head and shoulders" framing, where subjects' eyes are located about 1/3 of the way from the top of the screen. One common mistake that amateurs make is to frame the video subject in the center of the video. This makes the subject look short, with too much space above his head. The rule of thirds, shown in Figure 1, will suit you well.


Figure 1: Basic composition using the rule of thirds


To use the rule of thirds, divide your video image into thirds, both horizontally and vertically. You should try to place things of interest on the lines dividing the picture into thirds. Where the lines intersect are particularly good places. If you have a single subject, and you're shooting straight on, try to put the subject's eyes on the top 1/3 line. This makes for a well-balanced image that's pleasing to the eye. It doesn't matter how close or far away you are, the subject's eyes should remain on this line. If you get really close, you'll find that your shot crops off the top of his head (or maybe just his spiky hair). That's okay; if you place your subject at the center of the screen to try and keep his hair in the shot, it will look odd. Use the rule of thirds! It has been serving artists, photographers, and videographers for many years.

Another thing to consider is where the subject is looking. For a single talking head subject, it's best if they face directly towards the camera. In an interview situation, it's better if they are slightly to one side, looking toward where the other person is. In addition, it's most pleasing to the eye for the person to be angled toward the key light so that the part of the face getting the harsh light has the least exposure to the camera.

This may sound a bit complicated, but when you have your equipment set up it's easy to turn people one way or the other to see what effect it has on the shot. This is common practice in studios and is referred to as cheating. Sure, the subject may not be facing directly toward the interviewer, but if the shot looks better on camera, go with it.

Use a tripod
It is absolutely imperative to use a tripod when filming for the Web. Quite simply, using a tripod improves your video quality. Sure, most cameras come with built in handles that make them very portable, and carrying a tripod around is awkward and cumbersome. But the simple fact is that when you encode your podcast later, unnecessary motion will compromise your video quality, and hand-held content has lots of unnecessary motion.

Focus
It may seem obvious, particularly now that so many cameras have automatic focusing mechanisms built in, but it's critical that your subject remains in focus. Properly focused frames have more detail and consequently look better, even after encoding. Ironically, the auto-focus mechanisms of modern digital cameras can cause problems with your focus.

Auto-focus mechanisms work by making assumptions about what is most important in your frame. Things that are bright or moving tend to be interpreted as important. In many cases, this is fine, but if your subject is standing in front of a lake with boats sailing by, for example, the camera has a hard time deciding whether you're trying to shoot the static subject or the moving sailboats. Often the camera becomes confused and continually refocuses on different objects. In most situations, you're better off using the manual focus option if your camera offers one.

Focusing your camera manually is easy if you follow this simple procedure:

1. Zoom all the way in to your subject, and look at something with a lot of detail, such as the eyebrows or hairline.

2. Adjust the focus until it's as sharp as possible.

3. Zoom back out to your original shot composition.


That's it. Provided your subjects don't move too much, they'll stay in focus. If they do move, or if you decide to change the camera position, remember to re-focus each time.

White balancing and exposure
Earlier in this blog, we discussed color, in particular how our brains compensate for the differences in colors under different kinds of light. Cameras attempt to do this, but it's usually a better idea to manually white balance your camera to make sure your color representation is accurate. Manually white balancing a camera is a simple procedure. The idea is to "show" the camera what the color white looks like under the existing light. Given this information, the camera can then adjust its internal circuitry to compensate for the light, and the resulting video image will have faithful color representation.

Apr 20, 2008

Advanced EQ techniques

EQ also can be used as a corrective measure. The techniques described in the preceding sections aimed to improve the overall sonic quality of your audio by focusing on what you wanted to be heard. You also can use EQ to remove extraneous noise that you don't want in the file.

Clearing up noise
If you recorded your audio in a noisy environment, you can use EQ to get rid of some of the worst noise. First, roll off all the low frequencies. You should do this regularly to all frequencies below 60Hz, because they're generally not reproduced by most systems. You can roll off more if you're dealing with serious noise. For example, traffic outside a busy city window will be audible as a steady low rumble, with the occasional siren or horn. If you roll off more of the bottom end, the sound of the traffic will be less audible.

Similarly, you can roll off high frequencies if you have noise problems such as air conditioning or tape hiss. In the preceding example, we boosted all frequencies above 5KHz to give the file some air; clearly, this is not something we'd want to do if we were in a noisy environment. Of course, we could have used the shelf to roll off all high frequencies above 8KHz or so and added a slight lift around 4–5KHz. Very little important voice information is in the range above 10KHz, so if you have noise problems, you can safely roll this off without any damage to the intelligibility of the audio.

Of course, in extreme situations, you can get pretty savage with EQ if necessary. For example, sometimes news reporters outside during a storm sound like they're talking through a telephone line. This is because the audio engineer at the station has rolled off all low frequencies and high frequencies, leaving just the mids.

Pops
Sometimes, a pop sneaks into your audio file, even if you're using a pop screen. If you zoom in to your audio file and look at the offending pop, you'll see that it's a very brief, very loud low frequency burst. You can highlight the offending word (or even syllable) and roll off the offending bass frequencies, as shown in Figure 1.


Figure 1: A pop is visible as a short, loud low-frequency burst, which can be fixed with EQ.


Experiment with different amounts of roll-off at different frequencies. You'll find that if you roll off too much, you get rid of the pop, but the result may sound unnatural. Find a good balance between removing as much of the offending pop as possible, while retaining the natural feel of the original.

Dealing with sibilants
Some people have a problem with sibilants, which are consonants like the letter "s" ("d" and "t" can also be a problem but usually nowhere near as much) that contain a burst of high-frequency information. This is audible as a whistling or, well, an "ess" sound. A little of this is natural, but too much is annoying. Some people actually whistle when they pronounce these consonants, which is an extreme version of the same problem. Because this is usually a very localized problem, it can be dealt with using EQ.

Sibilants usually are most prominent around 5KHz. Just like the preceding example where a word was highlighted and bass frequencies were rolled off to get rid of a pop, you can highlight a word and try to cut the offending frequency. Dealing with sibilants is more troublesome than popping, because our ears are very sensitive to high frequencies, and even if it's only a momentary change, our ears may notice that the sound has changed. Not only that, but if someone has trouble with sibilants, it usually manifests itself throughout the entire interview — and the letter "s" is very popular in the English language. In general, if you want to deal with sibilants, you should use a de-esser. That's a fancy term we audio engineers made up to describe a "frequency dependent, side chain controlled compressor." De-essers are explained in the compression section, which comes next.