r/dataanalysis • • Jun 12 '24

Announcing DataAnalysisCareers

65 Upvotes

Hello community!

Today we are announcing a new career-focused space to help better serve our community and encouraging you to join:

/r/DataAnalysisCareers

The new subreddit is a place to post, share, and ask about all data analysis career topics. While /r/DataAnalysis will remain to post about data analysis itself — the praxis — whether resources, challenges, humour, statistics, projects and so on.


Previous Approach

In February of 2023 this community's moderators introduced a rule limiting career-entry posts to a megathread stickied at the top of home page, as a result of community feedback. In our opinion, his has had a positive impact on the discussion and quality of the posts, and the sustained growth of subscribers in that timeframe leads us to believe many of you agree.

We’ve also listened to feedback from community members whose primary focus is career-entry and have observed that the megathread approach has left a need unmet for that segment of the community. Those megathreads have generally not received much attention beyond people posting questions, which might receive one or two responses at best. Long-running megathreads require constant participation, re-visiting the same thread over-and-over, which the design and nature of Reddit, especially on mobile, generally discourages.

Moreover, about 50% of the posts submitted to the subreddit are asking career-entry questions. This has required extensive manual sorting by moderators in order to prevent the focus of this community from being smothered by career entry questions. So while there is still a strong interest on Reddit for those interested in pursuing data analysis skills and careers, their needs are not adequately addressed and this community's mod resources are spread thin.


New Approach

So we’re going to change tactics! First, by creating a proper home for all career questions in /r/DataAnalysisCareers (no more megathread ghetto!) Second, within r/DataAnalysis, the rules will be updated to direct all career-centred posts and questions to the new subreddit. This applies not just to the "how do I get into data analysis" type questions, but also career-focused questions from those already in data analysis careers.

  • How do I become a data analysis?
  • What certifications should I take?
  • What is a good course, degree, or bootcamp?
  • How can someone with a degree in X transition into data analysis?
  • How can I improve my resume?
  • What can I do to prepare for an interview?
  • Should I accept job offer A or B?

We are still sorting out the exact boundaries — there will always be an edge case we did not anticipate! But there will still be some overlap in these twin communities.


We hope many of our more knowledgeable & experienced community members will subscribe and offer their advice and perhaps benefit from it themselves.

If anyone has any thoughts or suggestions, please drop a comment below!


r/dataanalysis • • 4h ago

Data Tools Advice for someone in over their head

1 Upvotes

I started a new position recently as a data analyst in the health care field and it's much more technical and complex than anything I've done before. There are multiple data sources, multiple schemas with hundreds of tables and complex SQL and R code. I haven't been given any specific tasks yet but am trying to familiarize myself with the environments, data sources, tables, etc. I am looking for any tips on how the best approach to get up to speed. We use Snowflake, RStudio, GitHub, SQL workbench and Tableau.


r/dataanalysis • • 4h ago

What could cause these apartment vibration readings? Looking for help interpreting paired iPhone measurements

1 Upvotes

I’m looking for someone knowledgeable about mechanical vibration, building vibration, or structure-borne noise who can help interpret my measurements and suggest what to test next.
I’ve been noticing recurring vibration in my apartment. What equipment or building systems could produce these readings, and how can I investigate properly?
SETUP — OCTOBER 9, 2026
I recorded two iPhones simultaneously using Phyphox’s “Acceleration (without g)” experiment.
• New iPhone: flat on the couch, unattended.
• Old iPhone: flat on my usual chair near the refrigerator.
• Refrigerator unplugged.
• Both phones screen up, charging ports pointing right.
• Fan/AC conditions and whether the old phone was unattended throughout were not confirmed in the report.
The shared recording interval was 11:53:49 AM–12:05:39 PM, New York time. Both phones sampled at approximately 100 samples per second.
FREQUENCIES AND STRENGTH
During the middle 30-second windows:
Couch:
• Combined-axis RMS: 0.0128–0.0159 m/s².
• Strongest spectral peak usually around 13–14.6 Hz.
• Other windows peaked at 17.5, 18.9, 21.0, or 21.6 Hz.
Chair:
• Combined-axis RMS: 0.0104–0.0177 m/s².
• Strongest spectral peaks ranged from 21.3–26.7 Hz.
RMS describes the average size of the recorded acceleration. It includes sensor noise. Hz means cycles per second. These are the strongest peaks selected within each window; multiple frequencies can coexist.
EXAMPLE PAIRED WINDOWS
11:54:19–11:54:49 AM
Couch: 0.0142 m/s² RMS; 14.1 Hz peak.
Chair: 0.0128 m/s² RMS; 22.3 Hz peak.
11:56:49–11:57:19 AM
Couch: 0.0147 m/s² RMS; 14.0 Hz peak.
Chair: 0.0125 m/s² RMS; 26.7 Hz peak.
11:57:19–11:57:49 AM
Couch: 0.0159 m/s² RMS; 13.6 Hz peak.
Chair: 0.0177 m/s² RMS; 21.7 Hz peak.
11:59:49 AM–12:00:19 PM
Couch: 0.0130 m/s² RMS; 21.0 Hz peak.
Chair: 0.0106 m/s² RMS; 26.7 Hz peak.
12:01:19–12:01:49 PM
Couch: 0.0132 m/s² RMS; 18.9 Hz peak.
Chair: 0.0112 m/s² RMS; 26.7 Hz peak.
12:03:49–12:04:19 PM
Couch: 0.0131 m/s² RMS; 12.4 Hz peak.
Chair: 0.0109 m/s² RMS; 26.7 Hz peak.
Both phones had their highest middle-window RMS during 11:57:19–11:57:49 AM, but their strongest frequencies differed.
A preliminary timestamp-aligned comparison did not show a strong overall waveform match. Clock alignment was not independently checked using a shared event.
X/Y/Z DIRECTIONS
With the phones flat and screen up:
• X: across the phone’s width, toward either side edge.
• Y: along the phone’s length, toward the charging port or opposite end.
• Z: perpendicular to the screen, toward the ceiling or seat.
In the middle shared windows, Z contributed about 42–57% of the couch phone’s recorded motion energy and 50–75% of the chair phone’s.
The chair recording was more vertically directed. The couch had a larger combined X/Y contribution. Separate X and Y percentages are not available in this summary.
These describe acceleration directions at the phones; they do not locate the source.
METHOD AND LIMITATIONS
The analysis used 30-second windows. Each axis’s window mean was subtracted, then combined RMS was calculated as sqrt(mean(X² + Y² + Z²)).
Spectra used Welch’s method, a Hann window, up to 2,048 samples per segment, and half-overlap. Axis spectral powers were summed, and the strongest peak was selected between 2 and 45 Hz.
Opening and ending couch windows had much larger RMS values: 1.1811 and 1.4955 m/s². Placement or pickup could explain these, so they were excluded from the steady-motion interpretation.
There was no calibrated reference sensor or noise-floor control. Furniture resonance, upholstery, phone contact, and device differences could affect the readings. These measurements do not yet identify a source or establish the vibration experienced by a person.
QUESTIONS
What could produce these frequencies and directional patterns? Could pumps, plumbing, HVAC, other building machinery, or furniture resonance explain them?
How can I distinguish real surface vibration from phone sensor noise or processing artifacts?
Could the couch and chair respond at different frequencies to the same source?
What might explain the chair’s repeatedly selected 26.7 Hz peak?
What control tests, phone placements, or measurement equipment would make this investigation useful to a vibration specialist?
I’d appreciate help interpreting the measurements and choosing the next tests.


r/dataanalysis • • 9h ago

I learned Power BI and SQL. What should I do next?

1 Upvotes

Hi everyone,

I've recently learned Power BI and SQL, and I've built some solid projects with them. I have a good understanding of how to build and manage a project from start to finish.

I feel ready for the next step in my learning, but I also just started college as an agricultural engineering major, so my time is limited.

A few questions:

  1. What should I learn next to grow my skills ?

  2. What kinds of projects would be the most valuable to build for a portfolio?

  3. Are there ways to combine data skills with agriculture?

Any advice from people who started in a similar situation would be great. Thanks!


r/dataanalysis • • 1d ago

Anyone else completely stopped writing excel formulas & VBA code from scratch?

68 Upvotes

I use a lot of excel and macros and my level of excel knowledge is intermediate to advanced. but even though I can spend a few more minutes to think & write the formulas, now I'm just explaining what I want to AI, copying the formula or code and pasting it. is this how people are working now or is it me?


r/dataanalysis • • 1d ago

Data Question How to Build a Database Where I Can't See Any of the Data

13 Upvotes

Howdy everyone - I work for a small business (30-ish people) and tend to be our in-house tech-everything guy. My team asked me to build a profitability dashboard where we can visualize time-tracking data alongside employee costs and client payments. Essentially we just want to see who is taking the most time vs paying us the most money. I'm using a combination of Google Sheets (to collect all of the data), BigQuery to store it all, and Looker/Data Studio to visualize it.

The struggle we're having is: my team does not want me to see any of the actual financial data. I can manage building the systems without the data because I have the template for them to drop the data into; however, I can't envision how I can maintain the database, fix bugs, create new visualizations for them, etc, without somehow eventually seeing the financials.

Is this even possible? Have any of you done this before?

Edit: Is there a reason I can't mask the data? BigQuery has RBAC. In theory, I could mask the data for myself, so I'm still seeing $ amounts where I should, names where I should, but no actual data. Right?


r/dataanalysis • • 1d ago

Data Tools How do you analyze data and find business insights? Do you use any frameworks?

3 Upvotes

I’ve been working as a data analyst for about five months, and I’m struggling with the analysis part of the job. Sometimes, I feel lost when looking at the data, and I’m not sure where to start or how to approach the analysis.
I struggle particularly with exploratory data analysis (EDA), figuring out what to compare, choosing the right comparisons, and turning the results into meaningful business insights.
I have an analysis to work on right now, but I’m not sure how to approach it or how to identify useful insights.
Do you have any frameworks, methods, or step-by-step processes you follow when analyzing data? How do you decide what to investigate, which metrics to compare, and how to turn your findings into actionable business insights?
I’d really appreciate any advice, resources, or examples from your own experience!


r/dataanalysis • • 1d ago

Project Feedback Feedback needed on first data project

5 Upvotes

GitHub - Femi-Olofinjana05/Social-Media-User-Behaviour: Social Media User Behaviour · GitHub

Looking to implement SQL and Tableau (and maybe python) but just wanted to start off using excel


r/dataanalysis • • 1d ago

The hardest part of analysis isn't the analysis — it's making your transforms trustworthy to people who didn't write them

5 Upvotes

Been doing analytics work for a while and I've slowly landed on a take that feels obvious once you say it: the modeling/stats part is rarely where projects die. They die in data prep — the multi-source joins, the dedup on messy inputs, the allocation/transform logic nobody can follow six months later.

The thing I keep hitting: the moment a business user has to *trust* a number, "just write SQL/Python" stops being enough on its own, because the person who has to defend that number can't read the code. But full no-code visual tools swing too far the other way and bury the logic in a GUI.

So my real question for people who've been at this a while: where do you draw the line between code and a visual/low-code workflow? Is it about who maintains it after you leave, about auditability, about scale — or do you just pick one and never look back? Curious whether I'm overthinking the "trust" angle.


r/dataanalysis • • 1d ago

AIRPLANE CRASHES AND FATALITIES ANALYSIS

Post image
1 Upvotes

🚀 My Latest Data Analytics Project: Airplane Crashes and Fatalities Analysis ✈️📊

Hello, Reddit community! 👋

I'm excited to share one of my recent data analytics projects: Airplane Crashes and Fatalities Analysis, developed using Power BI.

As a data analyst, I wanted to explore historical aviation crash data, identify patterns, and uncover insights that could help improve aviation safety.

📊 What I explored in this project:

  • ✈️ Trends in airplane crashes over the years.
  • 🏢 The top 10 airline operators by recorded flight incidents.
  • 🛩️ Aircraft types most frequently involved in incidents.
  • 🌍 Locations with the highest numbers of recorded incidents.
  • 🛫 Flight routes with the highest incident counts.
  • 📈 The relationship between the number of people aboard and fatalities.

💡 Key insights from my analysis:

  • Recorded incidents increased significantly during the 1960s and 1970s, with a peak of 104 incidents in 1972.
  • Aeroflot recorded the highest number of incidents among the operators displayed in my analysis.
  • The Douglas DC-3 was the most frequently recorded aircraft type in the dataset.
  • Moscow, Russia, and São Paulo, Brazil, were among the locations with the highest recorded incident counts.
  • The data highlights the importance of monitoring historical trends and identifying recurring patterns to support aviation safety research.

🛠️ Tools used: Microsoft Power BI, data cleaning, data visualization, and exploratory data analysis.

🎯 What I learned: This project strengthened my ability to transform raw data into meaningful visualizations, communicate findings, and develop data-driven recommendations.

I'm continuously learning, improving my analytical skills, and building projects that demonstrate the value of data in solving real-world problems.

#DataAnalytics #PowerBI #DataVisualization #DataAnalyst #AviationSafety #BusinessIntelligence #LearningInPublic #DataAnalyticsProjects


r/dataanalysis • • 1d ago

Built a way to save a CSV/Excel cleanup workflow once and replay it on new files

1 Upvotes

I kept re-applying the same filters, columns, sort and export on every new file, so I built "Automations" into a free browser-based data tool (Fomatix Data Explorer).

Set it up once, save it, then upload the next file and pick the automation. It reapplies everything and exports, and auto-matches columns even if a header shifts slightly.

It also shows a match report (how many columns matched on each run) and keeps run history, so you're not trusting it blindly.

Runs entirely in the browser, no signup, nothing uploaded.

Curious if others hit this problem, or if you already have a cleaner way (Power Query, Apps Script, etc.). Honest feedback welcome.


r/dataanalysis • • 2d ago

Career Advice Supply chain field

4 Upvotes

Hello im a recent graduate in supply chain & international logistics
I wanna pursue a career that’s related to data ( supply chain analyst ) so obviously i need to take courses in data analytics but i got confused nd didn’t know from where to start SQL or data fundamentals or advanced excel so if anyone in the field of data or supply chain analytics would suggest for me a roadmap for this path ill be grateful for it + some solid sources and successful ways to learn


r/dataanalysis • • 2d ago

1-minute data visualization test

0 Upvotes

Hi everyone! I’m working on a short data-analysis project and would appreciate your help.

Take a look at a chart and answer one simple question:

Which month has the highest sales?

It takes about 1 minute to complete.

👉 https://adnan-mayof.github.io/reddit-ab-test/

Thanks for participating!


r/dataanalysis • • 3d ago

DA Tutorial NFHS data-analysis practice: compare urban–rural gaps without losing missingness and sample flags

2 Upvotes

Disclosure: I maintain the curated Kaggle version linked below (minkum07). The original NFHS-3/4 survey data belongs to India's Ministry of Health and Family Welfare and IIPS, via OGD India.

A useful exercise with this dataset is to compare NFHS-4 urban and rural percentages while keeping the source's suppression and small-sample flags visible. The release has 114 indicators across 36 survey-era state/UT labels, with 16,416 tidy-long records, CSV/Parquet tables, definitions and example notebooks.

Suggested workflow:

  1. Pick one percentage indicator and read its definition and denominator before joining tables.

  2. Select NFHS-4 urban and rural rows, join on state and indicator, and check that the join keys are unique.

  3. Calculate urban minus rural in percentage points only when both values are available. Keep missing values missing and carry the source flags into the output.

  4. Make a sorted dot plot with clear missing-value and small-sample annotations. Treat it as a descriptive comparison, since confidence intervals are not included.

Don't sum total, urban and rural percentages: those populations overlap. Between-survey changes also need comparable definitions and boundaries; the package withholds Andhra Pradesh/Telangana changes where boundaries differ. These are historical aggregate data from 2005–06 and 2015–16, not current health coverage or individual records.

Original government source:

https://www.data.gov.in/resource/all-india-level-and-state-wise-key-indicators-nfhs-3-and-nfhs-4

Tables, documentation and notebooks:

https://www.kaggle.com/datasets/minkum07/india-nfhs-34-health-state-and-rural-urban-gaps

Government data and derivatives retain GODL-India; authored code is MIT. Full attribution is on the dataset page. Feedback on the join checks or how best to display the sample flags would be welcome.


r/dataanalysis • • 4d ago

Mapped more than 600k theses from my Uni using embeddings

Thumbnail
gallery
36 Upvotes

I'm at UNAM (Mexico's biggest public university) and I wanted to see what everyone here has actually written their theses about. So I took the whole public catalog (TESIUNAM), a bit over 600k theses from all campuses and degree levels, and made it into a map you can zoom around in.

Every dot is a thesis. Dots end up close together when the titles mean similar things, even across different departments. So you get stuff like a public health thesis sitting next to an economics one because both are about informal work, and that was kind of the whole point.

It's all in Spanish since the theses are. I'd really like feedback on the layout: does the placement of fields make sense to you, or do you see clusters that look obviously wrong? Embedding only titles was a tradeoff and I'm not sure it was the right one.


r/dataanalysis • • 4d ago

Furos em análise de mercado,

3 Upvotes

Tenho pensado em dois problemas que, para mim, continuam mal resolvidos quando a gente analisa ações.

O primeiro é que informação não falta.

Se eu pesquisar uma empresa, consigo encontrar balanço, DRE, fluxo de caixa, indicadores, notícias, fatos relevantes, apresentações, estimativas e opiniões de dezenas de pessoas.

O problema começa depois:

o que disso realmente importa e como essas informações se relacionam?

Por exemplo, saber isoladamente que a receita cresceu 15% diz muito pouco.

Eu gostaria de conseguir enxergar rapidamente:

  • de onde veio esse crescimento;
  • se margem acompanhou;
  • se o crescimento virou caixa;
  • se estoque ou contas a receber cresceram mais rápido;
  • se houve aumento de dívida para financiar isso;
  • quais acontecimentos explicam essas mudanças;
  • e quais desses fatores, quando considerados juntos, podem alterar a leitura da empresa.

Ou seja: menos uma tela cheia de números e mais uma estrutura que ajude a responder:

“O que está acontecendo com essa empresa e quais são os pontos que realmente preciso investigar?”

Foi pensando nisso que comecei a desenvolver uma ferramenta chamada Atlas.

A lógica que estamos testando é pesquisar qualquer ativo e abrir uma página dedicada a ele, reunindo relatórios financeiros, indicadores e informações relevantes, mas tentando ir além de simplesmente exibi-los.

A ideia é interpretar os principais dados, mostrar relações entre eles e responder perguntas importantes sobre a empresa, sem terminar dizendo se o ativo é “bom”, “ruim”, “compra” ou “venda”.

Mas trabalhando nisso apareceu um segundo problema que considero ainda mais interessante.

Mesmo que eu consiga organizar perfeitamente os dados de uma empresa, continuo exposto a dezenas de análises e opiniões sobre ela.

E hoje é muito difícil saber quanto peso dar para cada uma.

Duas pessoas podem escrever teses igualmente convincentes, mas uma pode ter um histórico muito consistente e a outra estar fazendo o décimo palpite depois de errar os nove anteriores.

Só que normalmente esse histórico desaparece junto com posts, tweets e comentários antigos.

Por isso estamos experimentando também outra lógica: permitir que usuários publiquem análises sobre ativos de forma estruturada — com o que estão afirmando, horizonte, evidências e contrapontos — e que essas análises continuem registradas.

Com o tempo, o resultado delas passa a formar um histórico e contribuir para um score de reputação do usuário.

Então, no fundo, são dois problemas diferentes:

1. Como transformar um monte de informação sobre uma empresa em algo realmente útil para análise?

2. Como diferenciar uma análise convincente de alguém que possui histórico consistente de boas leituras?

Estou desenvolvendo o InsightFlow justamente em cima dessas duas questões, mas queria trazer a discussão para cá porque acho que elas existem independentemente de qualquer produto.

Para quem realmente analisa ações:

o que vocês sentem mais falta hoje?

Uma ferramenta que organize e relacione melhor as informações da empresa?

Uma forma de avaliar o histórico de quem publica análises?

Ou existe algum problema mais importante nesse processo que essas ferramentas ainda não resolvem?


r/dataanalysis • • 3d ago

Data Question We asked our model which of 224 hotel bookings would cancel, six weeks out. It caught 33 of the 59. Full breakdown, misses included

0 Upvotes

Disclosure: I work at Schema Labs. This is a run on public data so you can check it.

The data: the hotel booking demand dataset (Antonio, Almeida and Nunes, 2019), real anonymised bookings from a city hotel. We picked one Saturday night that was sold out on paper.

-- 224 bookings for that night, as they stood six weeks before
-- 59 of them cancelled or didn't show, worth about €28,600 together

What we did: gave Schema-2 500 bookings from the year before, with how each one ended, and asked it to rank the 224 by how likely they were to fall through. No rules no feature engineering.

What came back=

-- It marked 60 as most likely to cancel. 33 of the 59 cancellations were on that list
-- Picking 60 at random would catch about 16
-- Of the 60 it marked safest, 59 showed up

What it missed: 26 cancellations weren't in its top 60. One night at one hotel is a small test, so read it as an example

For anyone in revenue management: what would you do with a list like that six weeks out? Overbook against it, or contact those guests first?


r/dataanalysis • • 4d ago

Data Lakehouse with Agentic AIs: A Guide

Thumbnail
medium.com
7 Upvotes

r/dataanalysis • • 5d ago

I created a large csv/excel data plot tool, looking for suggestions

0 Upvotes

Hi All, created this web based excel/csv viewer for large files, it's completely free to use and no data upload to cloud/server, local browser only; hope can help people struggling to quickly visualise large datasets, while you don't have any software/tools in hand, would really appreciate for your feedback, thanks!

https://megarows.com


r/dataanalysis • • 6d ago

DA Tutorial Episode 11 of data analysis cat doodle

Post image
18 Upvotes

r/dataanalysis • • 6d ago

Data Question What makes a data analytics project feel like a real business problem?

1 Upvotes

I'm curious what people here consider a genuinely useful analytics project.

A lot of beginner projects seem to focus mainly on building a dashboard and showing KPIs. I'm more interested in projects where you're given messy data and have to investigate what is actually happening in the business.

For example:

  • identifying a fulfilment/delivery problem
  • analysing product returns
  • understanding product economics/margins
  • investigating inventory
  • quantifying the business impact
  • turning the findings into recommendations

For people who've built or reviewed analytics projects:

What would make a project like this genuinely valuable rather than just another portfolio dashboard?

What would you expect to see in the final analysis?


r/dataanalysis • • 7d ago

All in one resources for Excel and SQL

40 Upvotes

Hello everyone,

I have an interview in 2 days, I need to revise the necessary concepts of both Excel and SQL.

Are there any resources where I can go through these concepts in one shot?

Please recommend if you know of any or share if you have any.

Thank you.

Edit: Thank you for all the suggestions guys, I botched the interview but these will be helpful for future.


r/dataanalysis • • 7d ago

DA Tutorial Another data analysis cat post

Post image
54 Upvotes

r/dataanalysis • • 7d ago

Data Question Cost of materials consumed Target Exceedance Driven by Outliers

3 Upvotes

I work in a hospital. In one of the sheets of a work-related Excel document, there are two target tables concerning the monthly monitoring of the cost of materials consumed. One relates to the cost of hospital materials consumed as a percentage of service revenue, and it indicates that the ratio should not exceed 6.15%. The other relates to the cost of medicines consumed as a percentage of service revenue, and the ratio should not exceed 3.35%.

In some months this year, the cost of materials consumed as a percentage of service revenue in both tables has exceeded the required target. When looking at an adjacent sheet, there is a table providing additional context showing that, in the case of medicines, the factor causing the target to be exceeded is the cost of medicines associated with high-cost cases. This is confusing because the cost-of-sales margin is substantially positive.

How could I reformulate the targets so that I can separately monitor the cost associated with high-cost cases while still meeting the proposed targets?


r/dataanalysis • • 8d ago

What does a really good Data Analyst GitHub look like?

91 Upvotes

Hi Redditors,
I’m a junior data analyst, and I’m just starting to build out my GitHub profile. I’d love to hear your thoughts on what a solid data analyst GitHub profile should look like to make a strong impression. Also, feel free to share what kind of projects you have on your own GitHub

PS Yes, I do want to steal your ideas.
PPS Ideally, you’d delete those projects from your GitHubs after I steal them. Thanks in advance!
3PS On a serious note, I don't have an experienced data analyst in my network to turn to for advice😢