Sep 2019 — Mar 2020 · Chicago, IL, United States
Collaborated with student research collective to collect, analyze, and compile data from Twitter for water-related natural disaster insights.
Data Acquisition & Preprocessing: Automated the collection of tweets using the Tweepy API, focusing on keywords related to water-related disasters. Employed text preprocessing techniques including tokenization, stop word removal, and stemming, and leveraged FastText to generate high-quality word embeddings from the preprocessed text data.
Model Training & Evaluation: Utilized the Support Vector Machine (SVM) algorithm within the scikit-learn library to discern tweets related to water-related disaster events such as flash floods, tsunamis, droughts, etc. Trained the SVM classifier on a labeled dataset of tweets, applying cross-validation to ensure model generalizability. The model's performance was evaluated using accuracy, precision, recall, and F1-score, achieving an impressive accuracy of 85%. Key improvements were using GridSearchCV for hyperparameter optimization (i.e., learning rate, kernel), which improved the accuracy by 7%.
Containerization: Streamlined workflow by leveraging Visual Studio IDE within a Docker Container, fostering a standardized Linux & Miniconda environment.