Project

Analysis for data from the City of Walnut Creek's permitting and entitlement process. Figuring out how to better allocate resources for the process based on conclusions drawn from that data.

Data Given

We were provided with data that included the following:

  • visitor count by day
  • inspection stop and requests
  • receipt counts
  • account balances per client

This data was already processed and had a few visualizations. Data we would have liked to have in order to create more meaningful visualization would have included raw data indexed by specific permit with information pertaining to each permit.

Process

We first attempted to clean up the given data

  • Clean up data format for loading into pandas (remove summary/total rows, combine information from different sheets, etc.)
  • Deduplicate dates
  • Remove all non-numerical values for receipts, number of visitors, etc.
  • Replace references to 'holiday' or 'closed' with zero (i.e. number of visitors will be zero when office is closed)

We then decided to create a fake data set in the raw data format we would have liked to received the data in. We have included a snippet of the data set format.

our_image

The creation of the dataset containing 10,000 permit requests was uniformly distributed among 3,500 customers, filing and completion data from January 1st 2011 to the present, 26 different neighborhoods A-Z, and 10 different permit types. Each value in the Permit ID column is distinct (numbered 0-9999).

We completed our data analysis by uploading the fake data set to Tableau to create a functioning dashboard that provided information that can be filtered by year, permit type, processing time, and individual customer info.

Future Endeavors

  • Make an easy to use UI for clerks to use to input data
  • Clean the current data set to match our preferred format
  • Include more meaningful visualizations on the dashboard
  • Dynamically update dashboard
Share this project:

Updates