Gerade angezeigt 1 - 3 von 3
  • Some of the metrics are blocked by your 
    Item-typ:Veröffentlichung,
    Task-adaptable, Pervasive Perception for Robots Performing Everyday Manipulation
    Intelligent robotic agents that help us in our day-to-day chores have been an aspiration of robotics researchers for decades. More than fifty years since the creation of the first intelligent mobile robotic agent, robots are still struggling to perform seemingly simple tasks, such as setting or cleaning a table. One of the reasons for this is that the unstructured environments these robots are expected to work in impose demanding requirements on a robota s perception system. Depending on the manipulation task the robot is required to execute, different parts of the environment need to be examined, the objects in it found and functional parts of these identified. This is a challenging task, since the visual appearance of the objects and the variety of scenes they are found in are large. This thesis proposes to treat robotic visual perception for everyday manipulation tasks as an open question-asnswering problem. To this end RoboSherlock, a framework for creating task-adaptable, pervasive perception systems is presented. Using the framework, robot perception is addressed from a systema s perspective and contributions to the state-of-the-art are proposed that introduce several enhancements which scale robot perception toward the needs of human-level manipulation. The contributions of the thesis center around task-adaptability and pervasiveness of perception systems. A perception task-language and a language interpreter that generates task-relevant perception plans is proposed. The task-language and task-interpreter leverage the power of knowledge representation and knowledge-based reasoning in order to enhance the question-answering capabilities of the system. Pervasiveness, a seamless integration of past, present and future percepts, is achieved through three main contributions: a novel way for recording, replaying and inspecting perceptual episodic memories, a new perception component that enables pervasive operation and maintains an object belief state and a novel prospection component that enables robots to relive their past experiences and anticipate possible future scenarios. The contributions are validated through several real world robotic experiments that demonstrate how the proposed system enhances robot perception.
    Dissertation
      1297  493
  • Some of the metrics are blocked by your 
    Item-typ:Veröffentlichung,
    Perception for imagination-enabled robots
    Recent advancements in robotics and computer vision have enhanced object recognition and control strategies. However, these developments do not fully tackle the challenges of autonomous manipulation in dynamic, unstructured environments like households. Current systems often rely on specialized algorithms for perception, which lack generalizability and fail to verify the plausibility of their results. This thesis proposes a comprehensive framework that enhances robotic perception and manipulation in dynamic, unstructured environments by integrating a photorealistic, physics-enabled game engine. The core contributions of this research are threefold. First, it presents a unified perception architecture that combines imagistic reasoning, process-level control, and perception task adaptation within a single system. This architecture enables robots to construct internal hypotheses, simulate expected sensor data, and verify perceptual results against rendered scenes, facilitating grounded and introspective perception in real-world tasks. Second, the thesis presents a game-engine-based belief representation, utilizing real-time simulation as an internal model of belief states to enable high-fidelity visual hypothesis generation. The simulated environment represents a dynamic world model, including the robot state, allowing the system to assess the plausibility of perceptual results and predict the visual consequences of actions. Lastly, Perception Pipeline Trees (PPTs) are introduced as a modular process model for adaptive perception execution. PPTs combine hierarchical execution with flexible control flow, supporting reactive switching, concurrent processing, and introspective verification. This model accommodates conventional vision tasks and imagistic reasoning processes within a unified representation. The framework demonstrates effectiveness in real-world applications, including household assistance scenarios where robots perform tasks such as recognizing and manipulating objects, as well as tracking and interacting with humans. By enabling robots to not only observe but also reason about their environment through simulation, this work advances task adaptability, perception accuracy, and reasoning capability, laying the foundation for the next generation of intelligent, imagination-enabled robots.
    Dissertation
      75  60
  • Some of the metrics are blocked by your 
    Item-typ:Veröffentlichung,
    Uncertainty driven pose estimation - for rigid objects in CNN-based pipelines
    Nowadays, image-based object recognition and pose estimation are highly active research areas due to their importance in robotic perception and interaction. While modern CNN-based pose estimators achieve great results, they lack transparency regarding the trustworthiness and precision of individual estimates. This lack of certainty inhibits further processing of the results and deters deployment in production environments due to reliability concerns. As an answer, this thesis proposes a fusion-based approach in which, due to a novel output architecture, the CNN self-estimates the amount of information obtained, resulting in individual 6D uncertainty estimates per 6D pose estimate. Specifically, the CNN predicts the observed object points pixel-wise, along with the precision in the image plane of those predictions. All such gathered perspective information is then fused (without linearization) into a single, globally valid 13 × 13-sized information matrix, which is then regressed to yield the six-dimensional result. This separation allows the CNN to operate solely in image space, whereas the conversion from 2D image space to 6D pose is solved analytically. Additionally, the intermediate result of the globally valid information matrix facilitates the fusion with auxiliary information, such as depth, stereo, and prior knowledge, with ease, as it is simply a 13 × 13 matrix addition. With this approach, the pose is regressed from a fusion of all available data, unlike the more ad hoc approach of combining estimates in postprocessing. Also, the CNN call is wholly unaffected by the addition of these supplemental data. An extensive evaluation of the proposed architecture on multiple benchmark datasets showcases meaningful uncertainty estimates while maintaining competitive pose performance. Also, it shows that adding auxiliary information can significantly improve pose performance, but always relative to the amount of new information gained while maintaining the quality of the estimated uncertainty.
    Dissertation
      18  36