Gerade angezeigt 1 - 10 von 14
  • Some of the metrics are blocked by your 
    Item-typ:Veröffentlichung,
    Perception for imagination-enabled robots
    Recent advancements in robotics and computer vision have enhanced object recognition and control strategies. However, these developments do not fully tackle the challenges of autonomous manipulation in dynamic, unstructured environments like households. Current systems often rely on specialized algorithms for perception, which lack generalizability and fail to verify the plausibility of their results. This thesis proposes a comprehensive framework that enhances robotic perception and manipulation in dynamic, unstructured environments by integrating a photorealistic, physics-enabled game engine. The core contributions of this research are threefold. First, it presents a unified perception architecture that combines imagistic reasoning, process-level control, and perception task adaptation within a single system. This architecture enables robots to construct internal hypotheses, simulate expected sensor data, and verify perceptual results against rendered scenes, facilitating grounded and introspective perception in real-world tasks. Second, the thesis presents a game-engine-based belief representation, utilizing real-time simulation as an internal model of belief states to enable high-fidelity visual hypothesis generation. The simulated environment represents a dynamic world model, including the robot state, allowing the system to assess the plausibility of perceptual results and predict the visual consequences of actions. Lastly, Perception Pipeline Trees (PPTs) are introduced as a modular process model for adaptive perception execution. PPTs combine hierarchical execution with flexible control flow, supporting reactive switching, concurrent processing, and introspective verification. This model accommodates conventional vision tasks and imagistic reasoning processes within a unified representation. The framework demonstrates effectiveness in real-world applications, including household assistance scenarios where robots perform tasks such as recognizing and manipulating objects, as well as tracking and interacting with humans. By enabling robots to not only observe but also reason about their environment through simulation, this work advances task adaptability, perception accuracy, and reasoning capability, laying the foundation for the next generation of intelligent, imagination-enabled robots.
    Dissertation
      77  62
  • Some of the metrics are blocked by your 
    Item-typ:Veröffentlichung,
    Transforming Web Knowledge into Actionable Knowledge Graphs for Robot Manipulation Tasks
    One of the visions in AI based robotics are household robots that can autonomously handle a variety of meal preparation tasks. Based on this scenario, we present a best practice tutorial on how to create actionable knowledge graphs that a robot can use for execution of task variations of cutting actions. We implemented a solution for this task that integrates all necessary software components in the framework of the robot control process. In the context of this tutorial, we focus on knowledge acquisition, knowledge representation and reasoning, and simulating robot action execution, bringing these components together into a learning environment that – in the extended version – introduces the whole control process of Cognitive Robotics. In particular, the Tutorial will detail necessary concepts a knowledge graph should include for robot action execution, how web knowledge can be automatically acquired for the domain of cutting fruits, and how the created knowledge graph can be used to let robots execute tasks like slicing a cucumber or quartering an apple. The learning environment follows an immersive approach, using a physics-based simulation environment for visualization purposes that helps to illustrate the concepts taught in the tutorial.
    Konferenzbeitrag
      29  16
  • Some of the metrics are blocked by your 
    Item-typ:Veröffentlichung,
    An Ontological Model of User Preferences
    The notion of preferences plays an important role in many disciplines including service robotics which is concerned with scenarios in which robots interact with humans. These interactions can be favored by robots taking human preferences into account. This raises the issue of how preferences should be represented to support such preference-aware decision making. Several formal accounts for a notion of preferences exist. However, these approaches fall short on defining the nature and structure of the options that a robot has in a given situation. In this work, we thus investigate a formal model of preferences where options are non-atomic entities that are defined by the complex situations they bring about.
    Konferenzbeitrag
      20  28
  • Some of the metrics are blocked by your 
    Item-typ:Veröffentlichung,
    Towards Reactive Robotics with a Pinch of Image-Schematic Reasoning
    Today’s robots do not possess a deep understanding of interactions between physical objects that is also available to their behavior generation modules and as such show brittle performance in realistic environments. While this suggests a robot would therefore need more knowledge, its decisions should not be too complicated to arrive at or else the robot risks losing track of what matters from its environment. Thus, we investigate a mix of reactive approaches to robotics and reasoning, and propose a simplified theory of typical changes between image schemas. We show how this theory could be integrated in a robot’s perception-action loop, and describe some examples of using this theory to infer actions and perception queries for various stages of a pouring task. We are integrating this inference procedure into a simulated robot, but this integration is yet to be completed and as such, future work.
    Konferenzbeitrag
      34  16
  • Some of the metrics are blocked by your 
    Item-typ:Veröffentlichung,
    Prospective perception through cognitive emulation for robot manipulation tasks: "Perceiving like humans do"
    This thesis argues that bottom-up theories of perception suffer from the high semantic entropy arising from the severe spatial, temporal, and informational limitations of sensory input during everyday manipulation tasks. However, by emulating the “dark matter” of perception — including intent, functionality, utility, causality, and physis — and integrating it with sparse sensory data, robotic perception can achieve a causal, transparent, and computationally efficient ability to anticipate and explain relevant events in such tasks. To this end, the thesis introduces Probabilistic Embodied Scene Grammars (PESGs) to formalize this perceptual “dark matter.” It also presents a generator and a parser to respectively anticipate and explain event-centric scenes. The approach is demonstrated in complex real-world scenarios, including household tasks such as pancake making in kitchen environments, shopping tasks in supermarkets, and sterility testing tasks in medical laboratories.
    Dissertation
      104  66
  • Some of the metrics are blocked by your 
    Item-typ:Veröffentlichung,
    A Plan Executive Architecture for Transferable Robot Behavior: Generalized Action Plans and Their Context-Adaptive Execution in Real-World Settings
    (2025-10-02) ; ;
    Kaelbling, Leslie Pack
    ;
    To enable the deployment of robots beyond constrained laboratory settings into open-world environments, such as homes, retail stores or agile factories, robot software must be capable of transferring to novel execution contexts with minimal reprogramming effort. While state-of-the-art systems demonstrate impressive capabilities in controlled settings, transferring them to new environments, applications and hardware platforms remains a labor-intensive process. This thesis addresses the challenge of transferability in autonomous mobile manipulation systems by proposing a novel plan executive architecture that enables the reuse of robot behavior specifications across diverse execution contexts. The central innovation lies in combining robot control programs written in an expressive robot programming language with a novel mixed symbolic–subsymbolic action representation, termed action designators. This pairing yields generalized action plans that effectively combine action control flow and contextual reasoning. A hierarchical generalized action plan for the mobile pick and place action category is developed as a case study. It demonstrates context-adaptive execution by dynamically grounding its action designators through context-specific action parameter inference. This inference is carried out via a modular infrastructure that allows to integrate alternative and / or complimenting parametrization engines: a geometric world-state based engine, a heuristics-based engine, an experience-based engine trained on execution logs and an observation-based engine that learns from human demonstrations in virtual reality. A rapid simulation step is used to validate inferred action parameters prior to real-world execution. The architecture further supports self-specialization by refining generalized plans via learning and template-based plan transformations, improving performance in specific contexts. The approach is implemented in a fully integrated robot system that includes motion control, perception and learning components. It is empirically validated across 40 simulated and six real-world execution contexts, involving five robot platforms and a variety of environments and applications. Experimental results support the thesis hypothesis that a single generalized plan can effectively transfer across a broad range of execution contexts with limited reprogramming. The plan executive satisfies key requirements for transferability, scalability, extensibility, reactivity, failure tolerance, self-improvement, usability and explainability, advancing the capability of autonomous robots to competently act in diverse, real-world contexts.
    Dissertation
      38  49
  • Some of the metrics are blocked by your 
    Item-typ:Veröffentlichung,
    Between Input and Output: The Importance of Modelling Transients in Meal Preparation Tasks
    We are moving closer to autonomous robots preparing meals. While restaurant robots in static environments already are successfully performing single actions like making pizza, the goal is to enablerobots to perform changing actions, in various environments and with any available object. Towards this goal, a methodology for creating actionable knowledge graphs that can be used to parameterise general action plans has been proposed. However, for extended failure handling towards fully automated action execution, we argue that transients need to be considered. A transient can be described as a transitory object in a task that is not the same as the input object anymore but not yet the output object of the task. For example, when pouring ingredients into a bowl to make the dough, the added ingredients form a mass of ingredients (here: a transient) that only becomes dough through mixing them. This work shows how transients can be modelled and how robots can integrate and possibly benefit from this modelling.
    Konferenzbeitrag
      43  16
  • Some of the metrics are blocked by your 
    Item-typ:Veröffentlichung,
    ProductKG: A Product Knowledge Graph for User Assistance in Daily Activities
    The Web offers plenty of product information that is valuable for supporting decision processes. Research on Web knowledge acquisition and the Semantic Web has led to the creation of many domain ontologies and Web applications. What still is lacking is a connection of such knowledge to the real world. If object information is linked to environment information, users can get better, more personalised support in their daily activities like shopping or cooking since this enables them to link information about leftover products in the fridge to recipe information or a health profile to products the user is looking at in the store. It has been shown that semantic Digital Twins can successfully link object to environment information that can be used by agents like smartphone or service robot. Such semantic Digital Twins can offer even more services to users if they are connected to product information from the Web. This work introduces ProductKG, an open-source product knowledge graph integrating modular product information from the Web as well as accurate environment information from a semantic Digital Twin that can be customised for different applications and used devices as an example knowledge graph for assisting users in daily activities. We describe the design process and modularity of the knowledge graph as well as example applications of it, including an Augmented Reality shopping assistant, a dietary recommender and a hands-free recipe application. The modular ontologies enable personalisation of applications as well as accessing object information in relation to the current environment. We evaluate the acceptance of one example application through a user study. ProductKG is publicly available and will be maintained and extended over time in order to facilitate various applications such as in the retail and household domain
    Konferenzbeitrag
      12  30
  • Some of the metrics are blocked by your 
    Item-typ:Veröffentlichung,
    Dynamic Action Selection Using Image Schema-based Reasoning for Robots
    Dealing with robotic actions in uncertain environments has been demonstrated to be hard. Many classic planning approaches to robotic action make the closed world assumption, rendering them inefficient for everyday household activities, as they function without generalizability to other contexts or the ability to deal with unexpected changes. In contrast, humans robustly execute underspecified instructions in unfamiliar environments. In this paper, we initiate our research program where we propose the use of functional relations in the form of image-schematic micro-theories, formally represented in ISL𝐹 𝑂𝐿, to enrich action descriptors with semantic components. It builds on the body of work in embodied cognition showing that human conceptualization of action sequences is founded on abstract patterns learned from physical experiences in the form of spatiotemporal relationships between object, agents and environments. These theories are used to inform action selection mechanisms for behavioral robotics written in EL++ and we argue how these micro-patterns can be applied in a more general way to deal with underspecified action commands and commonsense problem-solving.
    Konferenzbeitrag
      80  53
  • Some of the metrics are blocked by your 
    Item-typ:Veröffentlichung,
    Ontological Context for Gesture Interpretation
    This study explores gesture interpretation by utilizing an ontological context. The aim is to store gestures and scene context data in an ontology and use its knowledge graph to actuate the robot arm to perform sets of manipulation tasks used in various environments. The knowledge graph captures the relationships between gestures, objects in the scene, and the desired actions. By putting the ontological context into use, the system can understand the meaning behind the gestures and execute the appropriate actions. The paper focuses on the development of the ontology, including the creation of class properties and the embedding of gestures within the ontology. Additionally, the paper explores how the integration of specifying context interpretation from the ontology may look to enhance the interpretation of gestures. The proposed approach aims to provide more intuitive and adaptive gesture-based supervisory control of robots in general. We tested the proposed ontological system in several tests so that it may be used in our future applications.
    Konferenzbeitrag
      43  27