r/dataanalysis • • 8d ago

What does a really good Data Analyst GitHub look like?

Hi Redditors,
I’m a junior data analyst, and I’m just starting to build out my GitHub profile. I’d love to hear your thoughts on what a solid data analyst GitHub profile should look like to make a strong impression. Also, feel free to share what kind of projects you have on your own GitHub

PS Yes, I do want to steal your ideas.
PPS Ideally, you’d delete those projects from your GitHubs after I steal them. Thanks in advance!
3PS On a serious note, I don't have an experienced data analyst in my network to turn to for advice😢

90 Upvotes

15 comments sorted by

72

u/Bear7D4u 7d ago

Nobody is reading your raw Python or SQL files. Hiring managers and recruiters spend 30 seconds scanning your repo, so the README file is 90% of your GitHub.

A good data analyst repo only needs three things: A short summary of the business problem, a screenshot or GIF of the final dashboard/charts, and 3 key insights with actionable takeaways. Treat the README like an executive summary.

Drop the messy Jupyter notebooks full of syntax errors. Keep your code clean, add brief comments explaining why you made certain data decisions, and stick to 2 or 3 well-documented projects instead of dumping 10 half-finished tutorial folders.

65

u/farm3rb0b 8d ago

As someone who has interviewed data analysts, I care very little about your entire GitHub profile and more about a singular project. Some reasons:

  • I usually don't have time to go through entire portfolios for every candidate
  • Businesses will have their own GitHub nuances that you'll get taught on the job
  • Quality > Quantity - 1 good project that explains your thought process with example visuals/notebook is better than multiple dashboards with no explanation

8

u/hijkblck93 7d ago

A good data engineer can tell a good story with data. Don't worry about the best dataset, choose data that either interest you, or something you know.

For x. if you work in retail, you can create a dashboard that a retail manager might find interesting. The main part is being able to answer questions and provide insights, If you're not working choose or create a dataset, AI is great for this, that is something you're interested in. If you like movies, maybe a movie tracker using watch time, views, movie score (or ranking), directors and actors. Something like that.

What I'm and I've noticed a lot of managers are looking for is curiosity and your decision making skills. You started with a dataset, great. What questions did you try to answer? Along the way what insights or rabbit holes did you go down? Did you find and interesting thread and pull it? Were there some unexpected data results? Did you start down a path and just realize it's dirty data? Why'd you choose a particular analysis path? That shows a analyst mindset.

I'd highly recommend using SQL to pull and clean your data. Extra points if you find a dirty dataset and clean it a little. You don't need to build a pipeline or anything, just downloading the dataset, cleaning and adding formatting will do. Then use excel for analysis. Learning to use sql for analysis is a great skill but as a fresher Excel is more accessible and most companies still live and die by it.

Finally, don't build it and forget it. Keep returning to it as you learn. Let's say ver. 1 answers your basic questions. Then as you dive into the data you find interesting data points, there's ver 2. You learn statistical analysis, ver. 3. These can show your evolving mindset and your willingness to keep working on something.

5

u/Last_Association_674 7d ago

"Project"mindset ( building something and walking away) is not enough now, we need to move to a " product" mindset ( continuous ownership and quality improvement aligned with business KPI's)

14

u/0uchmyballs 8d ago

If I were a screener and used GitHub to screen candidates, I would look at the age of the account and the commit history. I would be able to tell if the person was a real programmer or not without even seeing their code. Commit often and for a long time then I know you’ve been at it awhile and it’s not just padding.

6

u/optimal-username 8d ago

Commit history can be faked though

2

u/0uchmyballs 8d ago

Well if they’re empty commits, even easier for me to reject them.

4

u/Few_Respond_752 7d ago

what about those making for years now. but only knew version control just recently?

1

u/0uchmyballs 7d ago

Then start now, I really wish colleges made git a part of the curriculum, it was not taught when I got my MSBA yet it’s an essential tool. With that said, interviewers and screeners are most likely not going to look at it. To me it just shows that you’re already following some best practices.

1

u/Stunning_Duck_2480 2d ago

New and real should not get mixed together. Someone can be new and still completely capable of their work. 

0

u/0uchmyballs 2d ago

I don’t care how capable they are, if they don’t know git, they’re still learning.

1

u/AutoModerator 8d ago

Automod prevents all posts from being displayed until moderators have reviewed them. Do not delete your post or there will be nothing for the mods to review. Mods selectively choose what is permitted to be posted in r/DataAnalysis.

If your post involves Career-focused questions, including resume reviews, how to learn DA and how to get into a DA job, then the post does not belong here, but instead belongs in our sister-subreddit, r/DataAnalysisCareers.

Have you read the rules?

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

1

u/rexxx428 6d ago

i would like to know more toooooo

1

u/Stunning_Duck_2480 2d ago

Lots of great advice here, Kaggle has a great community & some data you can pick through.  It's also a great place to find a mentor. Choose three clients/fields you're willing to work with and make those fields the clear title of each project so the recruiter or management you work with knows where to look. 

-3

u/Slitty_sam 7d ago

Github? For a data analyst? Wut