Gerade angezeigt 1 - 2 von 2
  • Some of the metrics are blocked by your 
    Item-typ:Veröffentlichung,
    OCR Report
    Many social science researchers face the challenge of dealing with textual data that is only available on actual paper or ill-scanned PDF files, and require knowledge of image processing techniques and optical character recognition (OCR) software to obtain satisfactory results to enable further automated text post-processing. Based on sample scans of researches at the Collaborative Research Center “Global Dynamics of Social Policy” (SFB 1342), we compare the results of several open-source and commercial tools available for OCR. We evaluate each tool’s performance across three tasks, namely extracting plain text, recognizing the text style and its structure (hOCR), and extracting tables focusing not only the ability to accurately retrieve data from each cell but also the ability to properly capture the table layout. In this report, we summarize our findings and give recommendations for consideration when planning OCR projects.
    Bericht
    Band:
      337  267
  • Some of the metrics are blocked by your 
    Item-typ:Veröffentlichung,
    Filling Data Gaps in the Measurement of Income Inequality. A Complete Dataset of National GINI Coefficients 1995-2019
    Income inequalities are a major societal challenge (Grusky 2018; Polacko 2021). Despite the criticism that is being expressed, the Gini coefficient - especially income based - remains the most important indicator for measuring the extent and development of income inequality within a country. Unfortunately, Gini coefficients based on comparable methodologies are only available to a very limited extent. The most comprehensive data set available with consistent definitions for net income is the WIID Gini. With around 900 data points, this data set covers only 22% of the possible country-year combinations for the selected sample of 160 countries between 1995 and 2019. We pursue two objectives: (1) to close existing data gaps through statistical imputation thereby creating a consistent and plausible dataset of Gini coefficients for 160 countries with over 1 Mio. inhabitants from 1995 to 2019 and (2) to identify the socioeconomic and political indicators that most strongly influence these imputations. To achieve this, missing data are estimated using a gradient boosting machine (GBM) drawing on over 1.400 socioeconomic and political indicators from the WeSIS database. With this novel dataset, we enable researchers to broaden their inquiry into causes and effects of socio-economic inequality on a formerly unachievable scale.
    Buch
    Band:
      66  41