Gerade angezeigt 1 - 3 von 3
  • Some of the metrics are blocked by your 
    Item-typ:Veröffentlichung,
    Speech separation for monolingual and multilingual cocktail party scenarios
    (2025-10-02) ;
    Li, Haizhou
    ;
    Nakamura, Satoshi
    ;
    Li, Haizhou
    ;
    In everyday speech communication, humans face noisy multi-speaker soundscapes with overlapping sound sources, typically described as the cocktail party problem. Without much effort, humans can focus their attention on a specific voice or sound source while fading out the remaining voices or sound sources, referred to as selective auditory attention. For the development of future hearing aids and sophisticated algorithms for speech-based human-computer interaction, it is of great interest to equip machines with the same ability. The research field dedicated to the development of machine learning algorithms that embed this human ability goes by the name of speech separation. The underlying scientific problem is to develop algorithms that work in the wide range of everyday cocktail party scenarios just as the selective auditory attention ability of humans. This cumulative dissertation concerns deep neural network based single-channel speech separation and addresses five scientific problems (P): Target speaker absence and monologues (P1) deals with attended speakers in cocktail party scenarios who stop speaking for a while or hold monologues. Reliable target speaker reference (P2) formulates methods to enhance the reliability and stability of speech separation algorithms controlled by brain signals. Speech mode generalization (P3) has the goal to develop algorithms that work for multiple speech modes, such as normal and whispered speech. Cross-language generalization (P4) aims to develop algorithms that work for several seen and unseen languages. The multilingual cocktail party problem (P5) considers a cocktail party problem in which multiple languages are spoken and develops language-based speech separation algorithms. The contributions of this dissertation to solving P1-P5 mark a step towards the overall goal of equipping machines with a human-like ability of selective auditory attention that works reliably in real-world cocktail party scenarios.
    Dissertation
      120  96
  • Some of the metrics are blocked by your 
    Item-typ:Veröffentlichung,
    EEG-based Auditory Attention Tracking in Dynamic Multi-Talker Scenes
    (2026-08-14) ; ;
    Nakamura, Satoshi
    ;
    Li, Haizhou
    ;
    Humans can selectively attend to one speaker in the presence of competing voices and background noise, a phenomenon known as the cocktail-party effect (Cherry, 1953; Bronkhorst, 2000). Translating this ability into assistive listening technology remains difficult because an audio device must infer which source the listener intends to follow, often under short decision windows, multiple competing talkers, reduced sensor layouts, and changing spatial configurations (O’Sullivan et al., 2015; Van Eyndhoven et al., 2017). EEG-based Auditory Attention Decoding (AAD) addresses this problem by estimating the listener’s attentional focus from neural activity and is therefore an important research direction for future neuro-steered listening systems (Crosse et al., 2016; Ciccarelli et al., 2019; Geirnaert, Vandecappelle, et al., 2021). In this dissertation, such systems are treated as the motivating application context. The work develops EEG-based attention decoding and tracking methods, but it does not claim a clinically validated hearing-aid system, an embedded commercial hearable prototype, or validation with hearing-impaired users wearing hearing devices. The scientific problem underlying this dissertation is how to move EEG-based AAD from controlled proof-of-concept settings toward deployment-relevant research conditions. Prior work has reported promising results, yet many studies still rely on relatively static scenes, permissive window-level evaluation, subject-specific calibration, scalp-mounted sensors, or discrete attended-speaker labels (Ivucic, Pahuja, et al., 2024; Debener et al., 2015; Bednar and Lalor, 2020). Practical neuro-steered listening research requires a broader formulation: decoders should work with short windows without overstating real-time performance, should be evaluated under leakage-controlled trial- and subject-independent protocols, should remain robust when sensor layouts are reduced toward ear-centered recordings, and should extend from static speaker selection to continuous tracking of attended spatial trajectories. This cumulative dissertation addresses this problem through five connected research directions. First, it investigates short-window and reduced-montage binary AAD (bAAD) using cross-hemispheric attention in XAnet (Pahuja et al., 2023a). Second, it studies leakage-controlled evaluation for windowed EEG decoding and shows that cross-validation and segmentation choices can bias reported AAD performance when temporal dependence is not respected (Ivucic, Pahuja, et al., 2024; Pahuja et al., 2024). Third, it improves subject-independent AAD through physiologically motivated data augmentation and geometry-aware graph modeling (Pahuja et al., 2023b; Ivucic, Pahuja, et al., 2025). Fourth, it extends AAD to wearable multi-talker decoding (mAAD) with ear-EEG, where graph and cross-attention models capture intra-ear and inter-ear structure, and where DemoSAA demonstrates how four-speaker ear-EEG decoding can be organized in an interactive research pipeline (Pahuja et al., 2025c; Pahuja et al., 2025a). Fifth, it reframes AAD in moving-speaker scenes as tracking AAD (tAAD), where the target is not only a discrete attended talker but a time-varying attended azimuth trajectory estimated under causal and leakage-controlled conditions (Pahuja et al., 2025d; Pahuja et al., 2025b). Together, the results show that progress in EEG-based AAD depends on the joint design of model architecture, sensing strategy, task formulation, and evaluation protocol. Cross-attention and graph-based inductive bias improve short-window robustness; blocked evaluation gives more credible estimates of trial and subject transfer; augmentation and geometry-aware learning reduce subject-transfer degradation; ear-EEG can support multi-class attention decoding despite sparse sensing; and temporal tracking models enable reconstruction of attended spatial trajectories from EEG in dynamic scenes. The overall contribution is a dissertation-level argument that reliable neuro-steered listening research cannot emerge from stronger neural networks alone, but requires structured modeling combined with leakage-free, deployment-aware evaluation.
    Dissertation
      46  14
  • Some of the metrics are blocked by your 
    Item-typ:Veröffentlichung,
    Personalizing Myoelectric Silent Speech Interfaces via Cross-Speaker Training and Voice Timbre Control
    (2026-01-30) ; ;
    Nakamura, Satoshi
    ;
    Electromyography (EMG) signals, measuring muscle activity, are investigated for Silent Speech Interfaces (SSIs) to enable speech communication via silent articulation. The previous paradigm for EMG-to-Speech conversion relies on speaker-dependent models predicting acoustic features of speech from the same speaker providing EMG inputs. However, this approach makes SSI applications limited, as it 1) cannot be used to synthesize the personal voice of individuals unable to produce audible speech during EMG recording, 2) suffers from data scarcity, requiring each speaker to record a sizable corpus, and 3) leads to unintelligible speech in low-latency settings. The problem of converting EMG signals to speech in personal voices (1) is addressed by using voice conversion methods that disentangle phonetic and voice timbre information. The proposed voice-adaptive EMG-to-Speech models predict speech content features, mostly reflecting phonetic content, from EMG signals and combine them with reference audio of the target voice for speech synthesis. Further evaluations demonstrate that such models can be trained using EMG signals of silent speech only. The data scarcity problem (2) is addressed by several studies. For this purpose, EMG models are pre-trained with other biosignals, unlabeled EMG signals, and labeled EMG signals of multiple speakers, i.e., cross-speaker training. In particular, cross-speaker training improves average speech synthesis intelligibility, while eliminating the need to train speaker-specific models. To improve EMG-to-Speech in low-latency settings (3), this work presents an end-to-end model which outperforms previous low-latency baselines in speech intelligibility and naturalness while generating speech in less than 20 ms algorithmic latency. Furthermore, combining the previously outlined contributions, this work introduces a unified model which can convert EMG signals of multiple speakers to selectable voices.
    Dissertation
      87  69