# All posts tagged: World Bank

## Data Analysis and Interpretation Capstone

So, this is the end. It took six months, but today I completed and was certified for the Data Analysis and Interpretation Specialization by Wesleyan University through Coursera. When I first started in October 2015, I had no idea how to write code in Python, let alone produce graphs and run statistical analysis. It has been a fun experience learning how to write code in Python and learning the different kinds of statistical methods. Ironically, I learned these after I left graduate school. One would think that these are method courses you would take in school. For the Capstone Project, I do wish the data was more complete and over a longer period of time. It is difficult to run analysis on data that only goes back as far as 1972 and in many cases, missing records for many years in between. The results can be quite misleading, as it pointed to fertility rate as being highly correlated with environmental sustainability. However, fertility rate, in many cases is contingent on many different factors that are both quantitative …

## Capstone Project: Results

For those following my blog on my Data Analysis and Interpretation Specialization by Wesleyan University through Coursera, this is the final course and the Capstone project. Unlike previous courses, I will move away from urbanization data and try to tackle one of the problems provided by the course’s industry partner. This is my introduction. Below is our third assignment – the preliminary results. Results Only the results for Burundi, Ethiopia, and Liberia will be reported, as the other countries demonstrated no change or very slight change in the ensure environmental sustainability index. Descriptive Statistics: The following table shows the descriptive statistics for the Ensure Environmental Sustainability Index for each of the selected countries, starting from the lowest GDP per capita group to the highest. The standard deviations are much greater for the lowest GDP per capita group compared to the others. In three countries, Seychelles, Canada, and Ireland, no change in the value of the index was observed. It would appear that countries that reach a certain GDP per capita will have achieved a mean Ensure …

## Capstone Project: Methods

For those following my blog on my Data Analysis and Interpretation Specialization by Wesleyan University through Coursera, this is the final course and the Capstone project. Unlike previous courses, I will move away from urbanization data and try to tackle one of the problems provided by the course’s industry partner. This is my introduction. Below is our second assignment – the data management and analysis methods. Methods Sample: Out of the 211 World Bank recognized sovereignties, 8 (N=8) were chosen for this study. Countries that has the Ensure Environmental Sustainability goal were selected: three countries with the lowest GDP per capita (Burundi, Ethiopia, Liberia), three countries with the highest GDP per capita (Canada, Ireland, United States), and two from the median (Estonia, Seychelles). In addition to identifying associations between variables and the four sustainability indicators, this selection was used to also investigate how variable relationships differ in countries with varying degrees of economic development. Each country, depending on available data, has between 26 to 43 indicators for analysis with 36 years of data from 1972 to …

## Capstone: Variables Associated With Environmental Sustainability – A United Nations Millennium Development Goal

For those following my blog since the start of my Data Analysis and Interpretation Specialization by Wesleyan University through Coursera, this is the final course and the Capstone project. Unlike previous courses, I will move away from urbanization data and try to tackle one of the problems provided by the course’s industry partner. Below is our first assignment – the introduction to my final report. Variables Associated with Environmental Sustainability Using data provided by the World Bank, through DrivenData, this study looks to identify factors associated with the Environmental Sustainability Indicator defined as an United Nations Millennium Development Goal (MDG). Preliminary explanatory variables are Gross National Income, Forest Area, CO2 Emissions, Employment, Foreign Direct Investments, Household Final Consumption Expenditure, Adult Literacy Rate, Urban Population, Investments in Energy, and Energy Use. This mix of both economic and social factors will be examined for associations with the UN-MDG indicator of environmental sustainability. After the associated variables are identified, they will be used to create a model to predict data for the years 2008 and 2012. As a social/urban scientist interested …

## Chi-Square Testing…*Warning: It’s Painful*

Continuing with Data Analysis Tools… If you have not read my previous posts, I am currently enrolled in a Data Analysis Specialization with Wesleyan University through Coursera. With data from Gapminder, I am exploring a broad and basic question: does urbanization drive economic growth? For those of you interested in reading my literature review to gain a background on this project, please visit this page. For this assignment, I had to run Chi-Square tests on my variables. As always, both my Python and SAS codes are posted. Since all my data are quantitative, I had to first categorize them. Since I found a relationship between the absolute measure of urbanization (population in cities with over 1 million people) and GDP Growth rate, I decided to categorize GDP growth rate. Additionally, I wanted to see if there is a relationship between urbanization with the absolute measure of GDP  (GDP per capita). To categorize GDP per capita, I used cut-offs of 5000, 10000, and 100000 to produce three distinctive ranks whereby a country is poor if its GDP per …