Showing posts with label Journal club. Show all posts
Showing posts with label Journal club. Show all posts

Jun 29, 2011

2011/06/29


Nygaard, L.C., Herold, D.S., & Namy, L.L. (2009). The semantics of prosody: Acoustic and perceptual evidence of prosodic correlates to word meaning. Cognitive Science, 33(1), 127–146.

Presentation: Shelly
Summary: Thomas

As an attempt to blur the barrier between linguistic and non-linguistic aspects of spoken language, this study investigated whether prosodic correlates to meaning in different semantic domains could be reliably produced and perceived. In the production task, pictorial presentation of meanings from antonym pairs was used to elicit pronunciation of a given novel word. Acoustic analysis revealed that the overall valence of the meanings was reflected in the acoustic signal. Also, different antonym pairs seemed to elicit different acoustic features. The first perceptual task showed that listeners, upon hearing a novel word, could associate the word with the meaning by which the pronunciation was elicited in the production task. Yet it was possible that the results of Experiment 1 only revealed the overall valence in the acoustic signal, instead of domain-specific “semantics”. A follow-up perceptual experiment was done and the results revealed that the correspondence seemed to be domain-specific, because when a novel word for hot/cold was used to elicit responses to big/small, the meanings of the novel word were less successfully inferred than that from the same domain. The findings suggest that there are reliable and domain-specific prosodic correlates in speech.

Jun 22, 2011

2011/06/22


Wiget, L., White, L., Schuppler, B., Grenon, I., Rauch, O., & Mattys, S. L. (2010). How stable are acoustic metrics of contrastive speech rhythm? Journal of the Acoustical Society of America. 127(3): 1559–1569.

Presentation: Thomas
Summary: Sarah

This paper aimed to investigate the robustness of three proposed rhythmic metrics in the face of different sources of variability. The three metrics were %V (proportion of utterance comprised of vocalic intervals), VarcoV (rate-normalized standard deviation of vocalic interval duration), and nPVI-V (measure of durational variability between successive pairs of vocalic intervals). Factors tested in this study included speaker, material, and measurer. Five sentences were read by six speakers of Standard Southern British English. Measures were five human measurers and one automatic machine aligner. Results showed that among the three rhythmic metrics, nPVI-V was the most resistant to all sources of variability, whereas VarcoV was the least. In addition, it was found that material led to the largest variation, in which rhythmic scores fluctuated the most in different sentence conditions. Speaker and measurer produced the least variation, and high agreement was attained between the human measurers and the automatic aligner. It was therefore concluded by the authors that although the three metrics were generally robust measures, care should be taken, especially in the selection of sentence materials. It was recommended to always use the materials that accurately represent the phonological and metrical structure of the language.

Jun 15, 2011

2011/06/15


Wheeldon, L. & Waksler, R. (2004). Phonological underspecification and mapping mechanisms in the speech recognition lexicon. Brain and Language, 90, 401–412.

Presentation: Sarah
Summary: Sally

Different models have been proposed for recognizing phonological variations in speech. Previous studies have shown that some variations are tolerated in speech processing, while others are not. Controversy has been found between the possibility of underspecification in lexical representations and the nature of mapping mechanism. This study was designed to test which of the two hypotheses better accounts for the tolerance of phonological mismatch in English. In total, 100 stimuli (combinations of un-/changed and in-/appropriate for both unspecified and specified conditions, plus a control group) were embedded in sentences as testing materials. A pretest was given to ensure all the stimuli were read with equal clarity. For the experiment, accuracy and reaction time data were collected from a perception task using a cross-modal repetition priming paradigm. Participants were asked to make a forced-choice decision between changed and unchanged versions of the stimuli. Results showed that for underspecified stimuli, RTs in all primed conditions were faster than the control conditions, but no significant difference was found among different primed conditions (e.g. Unchanged In-/appropriate: They heard there was a wicked ghost/prince in the castle. Changed Inappropriate: They heard there was a wickib ghost/prince in the castle.). However, for specified stimuli, in addition to the same general pattern that all primed conditions were faster than the control conditions, the Unchanged Appropriate condition (e.g. She never had franctic moments with the twins.) was also responded significantly faster than the Changed Inappropriate condition (e.g. She never had franctip days with the twins.). In other words, different patterns were observed between underspecified and specified stimuli, as segmental change did not affect the degree of priming effect of the former, but resulted in different degrees of the effect to the latter. Since [+coronal] is believed a universal default place feature, these results supported the model with underspecified mental lexicon and a context-independent mapping.

Jun 1, 2011

2011/06/01


Van Engen, K. J. & Bradlow, A. R. (2007). Sentence recognition in native- and foreign-language multi-talker background noise. Journal of the Acoustical Society of America, 121(1), 519 – 526.

Presentation: Sally
Summary: Roger

Previous studies on speech-in-noise perception have shown that speech signal features have different resistance levels to degradation from noise. The present paper investigated how multi-talker babble, varying signal-to-noise ratios (SNRs), and different languages in the background noise affect the perception of English sentences. In the experiment, the target native-accented English sentences were embedded in English two- and six-talker babbles and similarly in Mandarin two- and six-talker babbles at SNRs of +5, 0, and -5 dB, which gave rise to 12 combinations (2 languages × 2 talker babbles × 3 SNRs). Sixty-six native English speakers were then asked to write down what they heard in the experiment. Results showed that overall higher SNRs yield better target sentence perception in all conditions. In agreement with previous studies, sentence perception is better in two-talker noise as opposed to six-talker noise. In term of the language of the noise, perception is better in Mandarin than in English noise in two-talker babble at SNRs of 0 and -5 dB. The first two findings are easy to interpret. The third finding has shown that the languages of interfering noise can affect the intelligibility of the target speech. The high density of noise in the six-talker babble seems to eliminate any information (linguistic) masking differences. In other words, the benefit of linguistic differentiation between the target and the noise which is accessible in the two-talker babble is overthrown in the six-talker babble. As for how linguistic differentiation facilitates the perception, several reasons, such as lexical effect, different phoneme inventories and syllable structures, prosodic factors, are suggested by the authors. However, exactly what factors or combinations of factors are taking effect remains undetermined and further researches are required.

May 25, 2011

2011/05/25


Francis, A. L., Ciocca, V., Wong, N. K. Y., Leung, W. H. Y., & Chu, P. C. Y. (2006). Extrinsic context affects perceptual normalization of lexical tone. Journal of Acoustic Society of America, 119, 1712–1726.

Presentation: Roger
Summary: Chris

This study investigated how listeners utilized contextual cues for the normalization of pitch range in tonal identification. According to the Pitch Range Assessment Model (PRAM), listeners usually normalize speakers’ pitch range based on contexts. The wider pitch range the context has, the better listeners’ normalization performance will be.  Experiment 1 tested how subjects normalized tonal range using two different types of contexts, one with a normal range of frequency variation and the other with the F0 held constant, which is equal to the average F0 of the context. Also, three kinds of degree shifts were manipulated to test the magnitude of contextual effect.  Results showed that shifting contextual pitch downward led to more high-level responses while shifting the pitch range upwards elicited more low-level responses.  Experiment 2 examined and compared the magnitude of contextual effects with the preceding and the following contexts. Results showed that raising the following context had a stronger effect than raising the preceding context. However, the effects of lowering the preceding and the following contexts were similar. It was suggested that the location of context should be taken into consideration in PRAM. Experiment 3 intended to answer the question whether the normalization of lexical tones was a linguistic process or merely an auditory process. The linguistic contexts were replaced with sounds generated by the hummed neutral vocal tract function in Praat. Results showed that listeners were unable to normalize lexical tones embedded in linguistically meaningless contexts. Experiment 4 examined how talker identity influenced tonal normalization. Results showed that no significant difference was found between the same talker condition and the different talker condition. In conclusion, the mechanism of estimating one’s tonal range is a process of extrapolation from the talker’s average F0. Furthermore, tonal normalization is a process involving integration of different positions (the preceding and the following contexts) and different sources of pitch perception.

May 18, 2011

2011/05/18


Aubanel, V. & Nguyen, N. (2010). Automatic recognition of regional phonological variation in conversational interaction. Speech Communication, 52, 577 – 586.

Presentation: Chris
Summary: Shelly

There is a growing interest in finding out the impact of dialectal variation on speech communication. The present study aimed to explore this issue by investigating the possible dialectal interaction in spontaneous conversion, with a focus on two major varieties of French, Northern French (NF) and Southern French (SF). An interactive task GMUP (Group’em up!) was designed, which is a collaborative game targeted to lead participants to spontaneously produce purpose-built names. In those names, five phonological dimensions were embedded, on which well-known differences between NF and SF exist: (1) word-final schwa realization in SF, (2) back mid vowel fronting in NF, (3) the contrast between mid-high and mid-low vowels in NF, (4) affrication of coronal stops in SF, and (5) the nasal vowel shift in NF. They recorded 12 dyads of one NF speaker and one SF speaker. The authors attempted to see whether the above five dimensions were distinctive enough for their Bayes classifier to determine the regional dialects of speakers, and whether speaker accommodation took place through out the game. That is, whether differences between NF and SF speakers would be mitigated throughout the game. Results showed that Byes classifier achieved a 79% recognition rate based on the five dimensions, suggesting that the developing classifier is quite good, and the above five dimensions are workable indicators on machines for differentiating the two dialects. However, such a high recognition rate did not decrease as the interaction between participants proceeded, meaning that speaker convergence did not happen. One possibility might be that the convergence took place at a subcategorical level, which their classifier was not sensitive enough to capture. Further studies will be needed to find out whether there are dialectal convergences throughout conversation at a more fine-grained level.

May 11, 2011

2011/05/11


Christophe, A., Gout, A., Peperkamp, S. & Morgan, J. (2003). Discovering words in the continuous speech stream: the role of prosody. Journal of Phonetics, 31, 585 – 598.

Presentation: Hsiao-chien
Summary: Shelly

Studies concerning word segmentation have found that allophonic cues, phonotactic rules, and stress patterns are used by both infants and adults for word boundary detection. However, whether cues from prosodic units play a role in this regard was less discussed. It was suggested that the coincidence of the target word boundary and an intonational phrase boundary can constrain the number of possible lexical activation triggered by the target word during on-line processing, which facilitates listeners to detect the target word at a faster speed. Therefore, this study intends to further the discussion on relevant issues by reviewing literatures exploring whether and how adults and infants benefit from prosodic boundary cues when trying to detect target words from sound sequences. Two experiments were reviewed. The first one was an experiment on French adults. They were asked to detect target words embedded in sentences as soon as possible. Results showed that for words which could be legally combined with the following syllable to form another words, detection was slower than those that could not. However, if an intonational phrase boundary was inserted inbetween the target word and the following syllable, the above differences between the two conditions disappeared and the listeners’ speed of detection became a lot faster. The results suggested that the existence of phrase boundaries indeed help on-line word segmentation. A similar experiment was done in another study on American infants, where word detection was observed through the head-turning paradigm. Results indicated the same patterns as those for the French adults in the first study. Results from the two studies suggested that phonological phrase boundaries are consistently used by human beings for word segmentation at a very early age. One of the possible causes may be that phonetic cues at phrase boundaries are usually strengthened, such as final lengthening, which amplifies the salience of the target words, resulting in faster detection by listeners. It was suggested that future studies would be needed to find out whether this is a universal pattern and what other boundary cues may be helpful in word segmentation, the results of which could be used to refine models for speech processing and language acquisition.

May 4, 2011

2011/05/04


Nakai, S. & Turk, A. (2011). Separability of prosodic phrase boundary and phonemic information. Journal of the Acoustical Society of America, 129, 966 – 976.

Presentation: Shelly
Summary: Thomas

Phonemic and prosodic information sometimes share the acoustic cues employed, as well as the time domain in which the cues are present. Such shared employment of acoustic cue and temporal domain requires the listener to decode the acoustic cues and retrieve information from both sources. The present study hypothesized that when prosodic information was encoded by multiple suprasegmental cues, the retrieval of phonemic information would be enhanced. Three sets of two-choice speeded and gated classification experiments were conducted to investigate the interactions between processing information on prosodic phrase boundaries and the place of articulation of stops. Experiment 1 showed that processing of phoneme was less intervened when there were multiple cues (duration and F0) to prosodic organization in the signal. Since the stimuli in Experiment 1 were not strictly controlled, two additional experiments were done to closely examine the effect of duration and F0. Results of these two experiments showed that when both duration and F0 cues were in the signal, the processing of phonemic information was less intervened than when only one of duration or F0 cues was present. The results were taken as evidence that integration of multiple prosodic cues can result in the relative ease of retrieval of both prosodic and phonemic information.

Apr 27, 2011

2011/04/27


Carlson, K., Frazier, L., & Clifton, C., Jr. (2009). How prosody constrains comprehension: A limited effect of prosodic packaging. Lingua, 119, 1066–1082.

Presentation: Thomas
Summary: Chris

This paper aimed to investigate how prosodic boundaries and prominence affect the resolution of sentence ambiguity. In Experiment 1, a two-choice sentence interpretation task was used to see the effect of prosodic boundary on matrix interpretation. Results showed that whether the sentence has one (ip-only) or two boundaries (ip and IP) does not affect sentence interpretation. Namely, the role for number of boundaries in interpreting replacives is small. Experiment 2(a) used sentences with semantic contexts that biased interpretation toward the subject in the matrix. Results showed that it was the accessibility of the antecedent of the replacives, rather than the prosodic boundary, that affected sentence interpretation, which failed to support the notion of prosodic packaging. In Experiment 2(b), disambiguation of subjects with animacy was required, which was to increase the bias toward the matrix. The percentage of matrix response was boosted but the effect of prosodic boundary was still absent. Experiment 3 varied the prominence of target words. Results showed that having an accent on the matrix subject increased the matrix interpretation. However, the reverse was not true for subordinate interpretation. In sum, lack of effects of prosodic boundaries on sentence interpretation/disambiguation suggests the limitation of prosodic packaging. Instead, prominence plays a crucial role in determining the accessibility.

Apr 20, 2011

2011/04/20


Snoeren, N. D., Segui, J., & Halle, P. A. (2008). Perceptual processing of partially and fully assimilated words in French. Journal of Experimental Psychology: Human Perception Performance, 34(1), 193–204.

Presentation: Sarah
Summary: Sally

Providing the first empirical data supporting the hypothesis that the role of context is modulated by assimilation strength, this study investigated the perceptual consequences of regressive voice assimilation in French. Pattern of this phonological variation is asymmetric, i.e., voiceless stops are usually fully assimilated in a voiced environment, while voiced ones are incompletely assimilated in a voiceless environment. Two experiments were conducted to test for perceptual compensation for assimilation in target words with voiceless and voiced stop offsets. Both experiments tested on native French speakers using an auditory-visual cross-modal form-priming paradigm. In Experiment 1, where the following context of the target words was absent, a stronger priming effect was found for canonical (unassimilated) forms than assimilated counterparts. In Experiment 2, in which the only difference from Experiment 1 was that the right context being made available, the same result was obtained. Furthermore, with the presence of the right context, the priming effect increased for assimilated voiceless-stop words, but no significant difference was found for assimilated voiced-stop words. It is suggested that for fully assimilated forms (the voiceless segments), the presence of the right context facilitates underlying form recovery, whereas this effect was not observed for incomplete assimilated forms (the voiced segments). Contextual information and bottom-up information, the two sources of information originated from assimilation, are thus believed to be complementary, and both are involved during the processing of assimilated forms.

Apr 13, 2011

2011/04/13


Flege, J. E., Takagi, N., & Mann, V. (1996). Lexical familiarity and English-language experience affect Japanese adults’ perception of /ɹ/ and /l/. Journal of the Acoustical Society of America, 99(2), 1161-1173.

Presentation: Sally
Summary: Roger

English /ɹ/ and /l/ are often misidentified by Japanese speakers who learn English in adulthood due to the fact that Japanese does not contain liquid consonants assembling English /ɹ/ or /l/. Japanese speakers’ misidentifications of English liquids might also be attributed partly to lexical factors. Previous studies found that inexperienced Japanese (IJ) subjects are more likely to misidentify liquids in a nonword that has a real word minimal pair than a real word that is minimally paired with another real word. One aim of this study was to determine whether experienced Japanese (EJ) subjects could identify word-initial English liquids at rates comparable to native English (NE) speakers. The other aim was to assess the effect of subjective lexical familiarity on the identification of liquids. Three groups (NE, EJ, and IJ) of subjects participated in two experiments. In the first experiment, 23 English minimal pairs, including both word-word pairs and word-nonword pairs (e.g. luck-ruck*) containing /ɹ/ or /l/ in the onset served as stimuli together with seven words beginning in /w/ and /d/ functioning as controls. Subjects rated familiarity of each word prior to the experiment and then identified onsets of these words presented auditorily. Results showed that while NE subjects identified both English liquids perfectly, EJ and IJ subjects obtained lower correct scores, with the EJ group higher than the IJ group. In addition, /l/ was misidentified more often than /ɹ/ by Japanese subjects. Plot of percent correct scores as a function of lexical familiarity also showed that familiar words were identified more correctly by both EJ and IJ groups. Comparing percent correct scores of NE and EJ subjects for three minimal pairs in which the two paired words had similar lexical familiarities suggested that EJ subjects were able to identify /ɹ/ but not /l/ at rates comparable to NE subjects. In the second experiment, stimuli consisted of eight additional minimal pairs made up of one word and one nonword together with 16 liquid portions edited out from these words. The order of the two conditions (whole-word and edited) was counterbalanced across subjects in each group, and 16 stimuli were presented randomly via headphone. Results showed that NE subjects identified liquids in the whole-word condition perfectly and only made two errors in the edited condition. For Japanese subjects, liquid tokens edited from nonwords had higher scores than the same liquid tokens presented in the whole-word condition. Lower scores in the whole-word condition were argued to be the result of a negative lexical bias associated with nonwords. However, patterns of real-word stimuli were different. Percent correct scores were higher for the /l/ tokens in the edited condition than in the whole-word condition. Yet scores for the /ɹ/ tokens in the two conditions did not show a significant difference. Signal detection theory was applied to explain these findings by assuming that /l/ had a broader normal sensory distribution than /ɹ/. Still, further research concerning the asymmetry between /ɹ/ and /l/ is required. Results also showed that, for /ɹ/ alone, scores for isolated tokens edited from words and nonwords were statistically the same for both NE and EJ groups, indicating again that EJ subjects were capable of identifying word-initial /ɹ/ so well as NE subjects. In summary, the present study shows the influence of lexical familiarity on the identification of English liquids /ɹ/ and /l/ which often sound ambiguous for native Japanese speakers. However, for /ɹ/, EJ speakers show ability to identify it at rates comparable to NE speakers when lexical factors are not relevant. Questions regarding the mechanisms of how lexical familiarity affects perception and how Japanese subjects’ perception of English liquids changes with more experience with English are yet to be answered satisfactorily.

Mar 16, 2011

2011/03/16


Kong, Y.-Y. & Zeng, F.-G. (2006). Temporal and spectral cues in Mandarin tone recognition. Journal of the Acoustical Society of America, 120(5), 2830–2840.

Presentation: Roger
Summary: Hsiao-chien


This study is aimed to evaluate the envelope and the fine structure cues used in Mandarin tone recognition for improving current cochlear-implant performances in pitch perception since both temporal and spectral fine structures are not explicitly encoded in the current processing schemes. The authors examined four types of acoustic cues in each experimental condition; first, for temporal envelope cues, the authors used the noise vocoder type of processing to manipulate the relative distribution of temporal and spectral information in the speech stimuli. Second, additional frequency modulation is used to produce better Mandarin tone recognition. Third, for presenting only the spectral envelope information in the absence of harmonicity cues, the authors made use of naturally recorded and LPC-synthesized whispered speech. Finally, the stimuli were manipulated by using the residue of a 14-order LPC processing.
Results of the first experiment showed that tone recognition performance was superior to the other three conditions in both quiet and noisy environments when both periodicity and spectral cues were available. Detailed spectral information (32 frequency bands) was required to produce tone recognition performance close to the original stimuli. Increasing the number of frequency bands improved tonal recognition, and frequency modulation (FM) contributed better tonal recognition. Tone recognition scores with spectral envelope cues were significantly poorer than those obtained with harmonicity cues alone in quiet and in all SNRs conditions. Overall, the present results are in agreement with previous studies of the effect of the number of bands. However, the authors found that temporal envelope cues are susceptible to noise, but spectral cues are more resistant to noise, which is considered a complementary contribution between temporal periodicity cues and spectral cues. This paper concluded that the fine structure had almost perfect performance than the envelope when observing the aspects of temporal and spectral cues.

Mar 9, 2011

2011/03/09


Gårding, E. (1987). Speech act and tonal pattern in Standard Chinese: Constancy and variation. Phonetica, 44, 13–29.

Presentation: Hsiao-chien
Summary: Sarah

This study intends to examine whether the intonation model proposed by the author for other languages also apply to a tone language such as Mandarin Chinese. Her model is composed of three essential elements, including turning points, pivots, and grids, all of which have distinctive functions. In particular, turning points are local F0 fluctuations that signal word boundaries; pivots refer to drastic changes of F0 direction, which correspond to syntactic boundaries; grids, delimited by pitch ceilings and pitch floors, serve the function of indicating different types of speech act. To testify the realizations of these three intonational parameters, a production experiment was conducted. Materials were four sentences with subject-predicate syntactic structures. Four native Mandarin speakers were recruited and asked to read each sentence in five speech act conditions (statement, yes/no question, statement after focus, statement with left focus, statement with right focus). There were three repetitions for each stimulus sentence. F0 patterns were tracked and compared across different speakers and conditions. Results showed that the four speakers realized tonal patterns similarly. That is, turning points were all anchored to segmental boundaries, and pivots all corresponded to syntactic boundaries. In addition, grids of different speech acts were also much alike across speakers, irrespective of the fact that grid width (pitch range) varied from speaker to speaker. As the results fit the model well, the author thus concluded that her model was applicable to the description of Mandarin intonation structures.

Mar 2, 2011

2011/03/02


Kirby, J. (2010). Dialect experience in Vietnamese tone perception. Journal of the Acoustical Society of America, 127(6), 3749–3757.

Presentation: Chris
Summary: Shelly

While speakers from different dialects may utilize different features for contrasting the same phonemes or tones, these differences do not seem to cause great impairment to mutual intelligibility between the dialects. This raises the question of whether or to what extent differences in production between dialects are indicative of differences in perception. The present study examined this issue by investigating the perception on tones in Northern Vietnamese (NVN) of both listeners from Northern and Southern Vietnamese (SVN). The cues to tonal contrasts in the two dialects have been shown to be different. NVN contains six tones, distinguished by F0 and voice quality, while SVN contains only five, distinguished just by F0. In the study of Brunelle (2009), it was found that such differences in dialect experiences indeed tune listeners to different cues, where NVN listeners are more sensitive to voice quality cues than are SVN listeners. However, since the paradigm used by Brunelle (2009) was an offline identification task, in which listeners might have enough time to access lexical information when making responses, the potential differences in prelinguistic processing might be obscured. Therefore, the present study employed a speeded online AX discrimination paradigm, attempting to assess listeners’ sensitivity to lower-level acoustic-phonetic information. Stimuli were syllables of different tonal pairs in NVN, on which listeners were asked to decide as soon as possible whether they were the same or different. Analyses on perceptual space and reaction time showed that tonal pairs with laryngealization were indeed much more confusable, and required more time to process for SVN listeners. Hierarchical clustering analyses further pointed out that SVN listeners tended to cluster laryngealized tones together. Nevertheless, it was also noted that, SVN listeners still showed some sensitivity to voice quality, an unfamiliar cue to SVN, when differentiating tones with and without it, but still, the saliency was lower than that to NVN listeners. Results reported in this study showed that listeners from different dialects are attuned to different cues. Therefore, the effect of dialect experience on tone processing at prelinguistic level can be confirmed.

Feb 23, 2011

2011/02/23


Ladd, D.R., Silverman, K., Tolkmitt, F., Bergmann, G., & Scherer, K.R. (1985). Evidence for the independent function of intonation contour type, voice quality, and f0 range in signalling speaker affect. Journal of the Acoustic Society of America: 78 (2), 435–444.

Presentation: Shelly
Summary: Thomas

Based on an earlier study (Scherer et al., 1984), there are two types of vocal cues to speaker affect. The continuous acoustic variables such as F0 range and voice quality reflect the states of the speaker in terms of physiological arousal, while the linguistic categorical variables such as contour types signal speaker’s attitude. The present study aimed to provide more evidence to these findings using three judgment experiments. Intonation contour type and F0 range were systematically controlled and varied. The first experiment was employed to see whether contour, range, and voice quality have independent effects in signaling affect and whether the first two variables are more related to speaker arousal than voice quality is. The materials contained three sentences (referred to as the text variable), which were resynthesized with two levels of contour, range, and voice quality, thereby generating 24 stimuli in total. The participants were asked to judge the affect of each stimulus on five bipolar 8-point scales for arousal and attitudinal variables, respectively. The result of ANOVA tests showed that range and voice quality had a strong effect on judgment related to speaker arousal; a wider range and harsh voice quality are signals of arousal, annoyance, and involvement. There were also effects of range and voice quality on cognitive attitudes, yet the effect sizes were smaller. On the other hand, contour, hypothesized to be more related to cognitive attitudes, had a significant effect on both arousal and attitudinal scales. As for interactions, no significant effect between contour, range, and voice quality was found. A follow-up experiment with a similar design was done to see whether the results on range and contour could be generalized when one additional factor, speaker, was involved. The other adjustment was that the ratings contained four 8-point unipolar scales for arousal and attitudinal variables, respectively. The results for range were replicated, yet only one significant effect of contour was found. Since only the effect of range seemed consistent, one additional RANGE x TEXT x SPEAKER experiment was done with five levels of range to see whether the effects of range were categorical or continuous. The results showed that the effect of range is linear for all scales, and the linear trend was highly significant, suggesting a continuous effect. To conclude, pervasive effects of range, especially on arousal scales, were found. As for voice quality and contour, the results were less consistent, which may be owing to the inability to systematically manipulate voice quality and insufficient understanding about the linguistic structure of intonation.

Feb 16, 2011

2011/02/16


Byrd, D., Krivokapic, J. & Lee, S (2006). How far, how long: On the temporal scope of prosodic boundary effects. Journal of Acoustical Society of America. 120(3), 1589–1599.

Presentation: Thomas
Summary: Sarah

While the lengthening effect at prosodic boundaries is extensively studied in both acoustic and articulatory domains, the scope of lengthening, nonetheless, is primarily investigated only in the acoustic, but rarely in the articulatory domain. In this regard, this study aims to explore the temporal scope of prosodic lengthening from the articulatory perspective, by tracking the movements of articulatory gestures at intonational phrase boundaries. In particular, both pre-boundary and post-boundary scopes were examined, as lengthening at these two locations are documented in the acoustic literature. Stimuli of the production experiment were three-syllable phrases all beginning with alveolar consonants (/d/ and /n/), placed before and after sentence boundaries. Articulatory sensors were adhered to subjects’ tongue tips. The measurements included the duration and displacement of the closing and opening gestural movements. Results showed that in the articulatory domain, the lengthening effect was only restricted to the immediate neighboring pre-boundary and post-boundary syllables. In addition, it was found that the opening phase of the pre-boundary syllable was more articulatorily strengthened, while for the post-boundary syllable, it was the closing phase that was more strengthened. Such findings therefore supported the π–gesture framework, proposed by Byrd and Saltzman (2003), which suggested that the boundary effect is anchored to the edge and diminishes as the syllables are further away from the boundary.

Jan 26, 2011

2011/01/26


Strange, W. (2010). Automatic selective perception (ASP) of first and second language speech: A working model. Journal of Phonetics. In press.

Presentation: Belinda
Summary: Sarah

In this paper, the author attempted to develop a perception model that describes and predicts the processing of speech signals by L1 listeners and L2 learners. In the Automatic Selective Perception model, native adult listeners’ perception of speech is considered as an automatic process. It is rapid and robust primarily because native listeners are able to selectively extract cues or parameters that are of contrastive salience. Native listeners are usually in the phonological mode when perceiving speech, in which low-level acoustic details are largely ignored. The opposite of the phonological mode is the phonetic mode. In particular, it is attention-demanding and usually requires listeners to pay attention to context-dependent acoustic information during speech processing. The phonetic mode is greatly involved in L2 perception, especially at the beginning stage of learning a second language. The objective of second language learning, therefore, is to obtain the selective perception routines, as called by the author, and to automatize them.
As these two modes are crucially relevant to L1 and L2 speech perception, it is necessary to study and discuss them separately. In the paper, it was shown that this can be achieved by manipulating stimulus complexity and task demands. The author discussed a series of experiments conducted in her laboratory, which looked into these two factors in detail. Results showed that stimulus and task manipulation yields different perceptual results of L2 listeners. Broadly speaking, L2 listeners would alter their modes of perception in face of different stimuli or tasks. These findings not only demonstrate her model’s validity, but also have important theoretical implications. The author suggested that more neurological evidence and training studies could be incorporated for further development of her perception model.

Jan 5, 2011

2011/01/05

Schmid, P. M. & Yeni-Komshian, G. H. (1999). The effects of speaker accent and target predictability on perception of mispronunciations. Journal of Speech, Language, and Hearing Research, 42, 56–64.

Presentation: Sarah
Summary: Sally

Intelligibility of nonnative speech has been investigated via different ways. Measurements based on accuracy of transcription by natives showed that accent ratings were not predictive of nonnative speakers’ intelligibility (Derwing & Munro, 1997; Munro & Derwing, 1995). Adapted from Cole’s method (1973 & 1980), the listening-for-mispronunciations task was first employed for nonnative speech in this study. During the task, mispronunciations (MPs) produced in the fluently articulated contexts were detected. Previous studies on native speech showed that MPs were more accurately detected when they were in predictable words, stressed syllables, and word-initial positions (Morton & Long, 1976; Cole & Jakimik, 1980; Cole et al., 1978).
In this study, 48 native listeners listened to native (N=4) and nonnative (N=4) production of English sentences. Among these sentences, target words were either of high or low predictability based on the context of the preceding words. All MPs were in word-initial position, with the target phonemes (6 stop sounds: /b/, /p/, /d/, /t/, /g/, /k/) manipulated for voicing, and place and manner (either fricative or nasal) of articulation. Results showed that the effect of sentence predictability was only significant in nonnative speech. Interaction between this factor and the effect of phonetic changes was also significant, as more MPs were detected in high-predictability sentences, especially for place and nasal changes. In addition, the effect of degree of accent was also observed: Native listeners were more accurate and faster in detecting MPs produced by native speakers. Based on a separate accent-rating task, the nonnative speakers were further divided into two groups; MPs produced by those with mild-to-moderate accents, were more accurately detected than MPs produced by those with a strong accent. Given that some sentences produced by nonnative speakers might not be equally comprehensible, and native listeners might adapt their expectations of acceptability in relation to speakers’ degree of accent, the authors concluded that intelligible but accented speech requires increased processing effort, which makes MPs not as detectable in nonnative speech.

Dec 22, 2010

2010/12/22

van Engen, K. J., Baese-Berk, M., Baker, R. E., Choi, A., Kim, M., & Bradlow, A. R. (2010). The wildcat corpus of native- and foreign-accented English: Communicative efficiency across conversational dyads with varying language alignment profiles. Language and Speech, 53(4), 510–540.

Presentation: Sally
Summary: Hsiao-chien

The present study was a part of the project of the Wildcat Corpus of native- and foreign-accented English in searching the further goal of creating a model of speech communication that integrates speech perception and production mechanisms with contact-induced sound changes. Using the framework of the corpus, the authors adopted two principles. “Talker-listener alignment” implies that speakers were from the same native language background. There are in total four types of conversations: native+native (N-N), non-native+non-native (NN1-NN1), native+non-native (N-NN), and non-native+non-native (NN1-NN2). Second, the corpus includes both scripted and spontaneous speech recordings. A new dialogue elicitation, the Diapix task, was developed. This task is a spot-the-difference game involving a pair of pictures and a pair of participants.
The authors recruited 24 native speakers of American English and 52 non-native speakers of English. The subjects were asked to cooperate with one other talker (of the same gender) in the Diapix task and then read a set of scripted English materials. In the Diapix task, all pairs of speakers were able to identify the differences, and the median score was 10 for all pair types. This indicates the meaningful comparisons of communicative efficiency across pair types. There were four sub-results: (1) The task completion showed that N-N pairs were faster than the other three pairs. The variance was much larger for the three groups involving NN talkers compared with the N-N group. (2) The analysis of balance of speech determined that N-N pairs were the least balanced group. However, the N partners spoke less than the NN partners in N-NN pairs. (3) The number of types in the conversations was similar among all pair types, but the word type-to-token ratios suggested that N-N pairs were more efficient than the other pairs. Moreover, compared to the native speakers in N-N pairs, natives in N-NN pairs had a greater amount of repetition in their interaction with non-native speakers. Non-natives in the N-NN pairs also had higher type-to-token ratios than non-natives in the NN1-NN2 condition. (4) The N-N pairs tended to proceed systematically, followed by N-NN pairs, and NN-NN pairs. Overall communicative efficiency was consistent with the pattern of alignment, in which N-N pairs were the most efficient, followed by N-NN pairs, NN1-NN1 pairs, and NN1-NN2 pairs.

Dec 15, 2010

2010/12/15

Cutler, A. & Chen, H.-C. (1997). Lexical tone in Cantonese spoken-word processing. Perception & Psychophysics. 59(2), 165–179.

Presentation: Hsiao-chien
Summary: Chris

This study aimed to explore the nature of perceptual processing of tonal information and to make a direct comparison with segmental perception. Speeded-response tasks were conducted, in which subjects had to make lexical decision on disyllabic Cantonese items. These items were constructed from real Cantonese words with the second syllable undergoing some alternation, be it onset, rhyme, or tone. Experiment 1 was designed to test which information was crucial to word recognition in Cantonese. Results showed that tonal mismatch elicited the highest error rate, followed by vowel mismatch. Syllables that differed in onset and tone were the easiest type. In the tone condition, there was a tendency for Tone 1 to have lower error rates compared to other tones. The goal of Experiment 2 was to assess the order in which perceptual information becomes available to native listeners and the speed with which a fairly distinct and a fairly non-distinct tonal difference can be perceived. A same-different judgment task was used. The test material contained two conditions: distinct (high-falling Tone 1 vs. mid-rising Tone 2) and non-distinct (low-falling Tone 4 vs. low-rising Tone 5), where distinctiveness depended on whether the tones of a pair had similar onset F0. Results showed that responses to stimuli that were in the distinct condition were significantly faster than those in the non-distinct condition. However, it was not sure whether the results had linguistic implication or were just acoustic in nature. Hence, Experiment 3 tested Dutch non-native speakers of Cantonese. Results were similar to Experiment 2. These results showed that in Cantonese, many tonal discriminations were hard to make, both for native speakers and non-native speakers. When distinctive tonal information occurs early, perceptual processing would be more efficient. However, tonal information does not usually arrive early since tones are primarily realized upon vowels and cannot be processed independently. Similar to the results in lexical stress languages, information is processed as soon as it becomes usable, but prosodic information may reach this state later than segments, which makes utilization of prosodic cues slower and more difficult than segmental cues in the process of word recognition.