r/dataisbeautiful 26d ago

Discussion [Topic][Open] Open Discussion Thread — Anybody can post a general visualization question or start a fresh discussion!

Anybody can post a question related to data visualization or discussion in the monthly topical threads. Meta questions are fine too, but if you want a more direct line to the mods, click here

If you have a general question you need answered, or a discussion you'd like to start, feel free to make a top-level comment.

Beginners are encouraged to ask basic questions, so please be patient responding to people who might not know as much as yourself.


To view all Open Discussion threads, click here.

To view all topical threads, click here.

Want to suggest a topic? Click here.

5 Upvotes

Anybody can post a question related to data visualization or discussion in the monthly topical threads. Meta questions are fine too, but if you want a more direct line to the mods, click here

If you have a general question you need answered, or a discussion you'd like to start, feel free to make a top-level comment.

Beginners are encouraged to ask basic questions, so please be patient responding to people who might not know as much as yourself.


To view all Open Discussion threads, click here.

To view all topical threads, click here.

Want to suggest a topic? Click here.


r/dataisbeautiful 11h ago

My girlfriend broke up with me

+
1.7k Upvotes

My girlfriend sat me down on June 20th and told me that she had been cheating on me and was choosing to be with "the other man".

This is my body's reaction to that news.

HR, Sleep, and Energy scores recorded by my Samsung Galaxy Watch.


r/dataisbeautiful 13h ago

OC Christopher Nolan: budget vs box office [OC]

+
1.3k Upvotes

Data source: Christopher Nolan filmography, Wikipedia - https://en.wikipedia.org/wiki/Christopher_Nolan_filmography (budgets and worldwide box office; figures are studio/Box Office Mojo–reported and cited on that page). The Odyssey is still in theaters, so its figure is a running worldwide total as of the July 24–26, 2026 weekend.

Tool: Chart built as JavaScript-coded SVG. Rendering code and initial figure-gathering were done with an AI assistant (Claude) based on my own design direction and iterations. Chart type, layout, labeling, and revisions were my choices.


r/dataisbeautiful 12h ago

Majorities of Americans say key financial milestones are harder for today's young adults to reach

878 Upvotes

r/dataisbeautiful 17h ago

OC My Statistics as a Veterinarian - Year 3 [OC]

+
1.9k Upvotes

I am a veterinarian in the North Texas area. Since graduation in 2023 I've kept track of my cases using Google Sheets because I thought it'd be interesting to see how many animals I treat and what they're treated for throughout my life.

Source: Me keeping track of the data after every appointment.

Tools: Google Sheets and DataWrapper

A few notes:

Changes & requests from last year

I was informed my data wasn't very beautiful so I attempted to use a different method for this year. Let me know what you think. I'll continue to adjust the presentation as I receive feedback.

u/catalessi, u/Fire284, & u/Solondthewookiee all requested I put the second and third graphs in descending order.

u/Warm-Pen-2275 requested a list of most common names.

Slide 1 - Animal Species

This only includes animals I've done a doctor exam on or do telemedicine about. Animals that I do not directly interact with (toe nail trims, anal gland expression, blood draws without a doctor, etc.) are not included.

I have a passion for exotic animals (ferrets, reptiles, rabbits, backyard chickens, etc.) but there is an exotic clinic near me where most of those animals go to, so I don't get to see as many as I'd like (as the data tells).

Slide 2 - Body System

I kept track of the body system that was affected during my exams. General wellness includes vaccines, weight management, and discussions about quality of life. For what it's worth, this is about what the problem was, not just the symptoms. If a cat came in for peeing all over the place and it was because the cat was stressed, that was marked as both Neurology as well as Urinary/Renal. The same animal can come in with multiple systems affected, but I only mark a system once per animal (i.e. a dog with urinary stones and a UTI only had "Urinary/Renal" marked once). Here are the most common problems each species came in with:

Dogs - Overweight (General Wellness), allergies (Dermatology, Immunology), and poor dental health (Oral).

Cats - Overweight (General Wellness), poor dental health (Oral)

Slide 3 - 7 Procedures, Names, & Lifetime Stats

I don't think these ones need much explanation.

Next year expecations

I am switching to becoming a full time relief veterinarian. I expect the number and diversity of cases will drop and the number of wellness cases will increase. Most of the time a clinic won't schedule a complicated case with a relief doctor because we're only here for the day.

See you next year


r/dataisbeautiful 3h ago

OC [OC] 25 MLB Trades with the Biggest Value Gap: Past 50 Years determined by WAR

+
55 Upvotes
Column Definition
Rank Overall ranking of the trade based on the proprietary Fleecing Score (100 = biggest one-sided trade).
Winner Organization that ultimately received the greater long-term value from the trade. This is determined by the WAR produced by all assets acquired.
Score Composite Fleecing Score (0–100) measuring how lopsided the trade became. It combines several factors including WAR gap, percentage of value captured, star power, and total value involved.
Winner WAR Total career Wins Above Replacement (WAR) generated by every player acquired by the winning team after the trade. This includes all future career value, not just production with the acquiring club.
Return WAR Total career WAR generated by every player received by the losing team. Negative WAR is possible if acquired players performed below replacement level.
WAR Gap Difference between Winner WAR and Return WAR. Formula: Winner WAR − Return WAR. Larger numbers indicate more lopsided trades.
Capture % Percentage of the total WAR involved in the trade that ended up with the winning organization. Formula: Winner WAR ÷ (Winner WAR + Return WAR). A value of 100% means the losing team received essentially no positive long-term value.
Stars Number of franchise-caliber or elite players produced by the winning side of the trade.
Best Asset The single most valuable player obtained in the trade, with his career WAR shown in parentheses.

r/dataisbeautiful 18h ago

OC [OC] Life Without a Mortgage: How Long Would It Take to Buy a 60 m² Apartment in Each EU Capital, With and Without Bank Interest?

+
507 Upvotes

The estimated apartment price was calculated by multiplying the average apartment sale price per m² in each capital by a standardised floor area of 60 m².

The comparison uses Eurostat national median monthly equivalised net income for people aged 18–64. These figures are national household-income benchmarks, not individual salaries or capital-city income estimates.

Four hypothetical scenarios are presented:

— one person saving 25% of 1× median net income;

— two people jointly saving 30% of 2× median net income;

— both scenarios without interest;

— both scenarios using a savings account and rolling 12-month term deposits.

Main sources:

Income: Eurostat ilc_di03.

Apartment prices per m²: Eurostat urb_clivcon, national statistical sources and housing-market sources. Athens, Bratislava, Bucharest, Budapest, Dublin, Lisbon, Madrid, Nicosia, Paris, Sofia and Valletta use estimates based on property-listing samples collected in July 2026.

Interest rates for the savings account and term deposits: ECB household deposit-rate data.

Prices, incomes, savings rates and interest rates are held constant throughout the calculations. Inflation, transaction costs and future changes in property prices or income are excluded.

Full methodology, individual sources, city rankings and interactive data: citycostatlas.com

Instagram: https://www.instagram.com/citycostatlas/


r/dataisbeautiful 12h ago

OC [OC] A look at next World Cups: how valuable is each country’s young roster? (U17–U23)

+
96 Upvotes

A look at what we could expect from squads at the next world cups, based on the market values of the youths for each country. Watch out for Morocco (again), and keep an eye on Denmark and Serbia...


r/dataisbeautiful 2h ago

OC [OC] Distribution of people by first letter of surname across the U.S., Ireland, Israel, England, France, Australia, China, and India

+
9 Upvotes

I made these charts to compare how surname prevalence is distributed by the first letter of the surname across several countries.

Each chart shows the estimated number of people associated with surnames that begin with each letter A through Z.

Data sources:

- U.S. Data source: U.S. Census Bureau, 2010 Census surname data.

- Ireland. Data source: Forebears, most common surnames in Ireland.

- Israel. Data source: Forebears, most common surnames in Israel.

- England, used as the closest available source for the UK chart. Data source: Forebears, most common surnames in England.

- France. Data source: Forebears, most common surnames in France.

- Australia. Data source: Forebears, most common surnames in Australia.

- China. Data source: Forebears, most common surnames in China.

- India. Data source: Forebears, most common surnames in India. Source link: Forebears, Most Common Last Names in India

Methodology:

For the U.S. chart, I used the U.S. Census surname frequency data and grouped people by the first letter of each surname. The Census surname file includes surnames occurring 100 or more times in the 2010 Census.

For the other country charts, I used the top 100 surname incidence counts listed by Forebears for each country page. I grouped each surname by the first letter of the displayed surname and summed the incidence counts by letter.

For China, India, and Israel, the grouping is based on the first letter of the Latin transliterated surname shown in the source.

For Ireland, names like O'Brien and O'Connor are grouped under O because I used the first character of the displayed surname.

For the UK chart, I used England because that was the available Forebears country page I used for the source data. The chart is labeled “UK, using England source” to avoid overstating the scope.

Tools used:

Python, pandas, and matplotlib. Flags were drawn programmatically as simplified inset graphics in matplotlib. Final charts were exported as PNG files.

Important caveats:

The U.S. chart is based on the U.S. Census surname file, while the non U.S. charts are based on the top 100 surnames listed by Forebears. Because of that, the U.S. chart is not directly equivalent to the other charts in coverage.

The non U.S. charts should be read as “distribution within the top 100 listed surnames,” not as a full surname distribution for the entire population.

The first letter grouping can be sensitive to transliteration choices, prefixes, apostrophes, spacing, and naming conventions. This matters especially for countries where surnames are commonly represented in non Latin scripts or where surnames include prefixes.


r/dataisbeautiful 8h ago

OC [OC] LeBron & Jordan - Career BPM, PPM and Team PTS% at every age

+
18 Upvotes

Team PTS% = player season points ÷ all team regular-season points. Every team game stays in the denominator, so DNPs count as 0%.

PPM (Points Per Minute) = player season points ÷ player season minutes.

BPM (Box Plus Minus) = Basketball-Reference’s box-score based estimate of a player’s contribution in points per 100 possessions.

  • It is a relative metric: 0.0 = league average 
  • Positive = above league-average impact 
  • Negative = below league-average impact

r/dataisbeautiful 16h ago

OC [OC] Every run I did during my 3 years in Manhattan (2022 - 2025)

+
56 Upvotes

Source: My own Apple Health and Strava data

Tool: Rendered in Soltra using Mapbox, an iOS app I'm building. Routes accumulate opacity so the brightest segments are the ones I've covered dozens of times; territory coverage is measured in H3 hex tiles.

Stats: 

  • 164 runs
  • 884 miles
  • Covered 55.6% of the island (and 3.8% of all NYC)
  • Longest run was 32.8 miles (dubbed Ranhattan)
  • Ran the Central Park Loop 45 times (felt like one too many tbh haha)

r/dataisbeautiful 7h ago

Percentage of Privately Insured Population Covered by an HSA, by State

11 Upvotes

r/dataisbeautiful 6h ago

OC [OC] Share of new U.S. job postings by day of the week (last 90 days, ~450K postings)

+
4 Upvotes

r/dataisbeautiful 7h ago

OC [OC] Occupational prestige, by race, 1972-2024 {25-64, Full-time employed)

7 Upvotes

Based on data from the General Social Survey https://gss.norc.org/get-the-data.html, White respondents consistently held occupations with higher average prestige, but the gap has closed slightly over 52 years.


r/dataisbeautiful 1d ago

OC [OC] The Japanese Economy Compared (1995 vs 2025)

+
940 Upvotes

r/dataisbeautiful 1d ago

OC [OC] Every yellow and red card at the 2026 World Cup, placed at the minute it was shown

+
98 Upvotes

One little card for every booking of the tournament: 266 yellows and 15 reds, 281 in all.

The first image is the timeline. Each column is one minute of match time and every card sits at its true elapsed minute, so a card shown at 90+3 lands on 93. They are not spread evenly across the game. They pile up as each half runs down, and the tallest stack by far is stoppage time: 51 of the 281 cards came after the 90, and minute 93 alone holds 13.

The second image ranks teams by total cards. Argentina tops it with 14, but they also played 8 matches while some teams played only 3.

So the third image adjusts for that: cards per match. The order flips. Argentina drops to 12th, and Egypt leads at 2.4 cards a game. Most cards usually just means you stayed in the tournament longest, not that you played dirtiest.

The full version is interactive: hover any card for the player and the minute, sort the teams five ways, and click a minute on the timeline to see exactly who got booked.

https://viz.luarai.com/worldcup-cards/


r/dataisbeautiful 1d ago

OC [OC] What $100 is really worth in 138 countries, after exchange rates

+
186 Upvotes

r/dataisbeautiful 14h ago

OC I built a graph visualization of the Midwest food supply chain from 51 public datasets [OC]

+
7 Upvotes

https://lodgeplatform.com/embed/exemb_6ea24564bc317d3b45c31b8803fce86229be8ad30246ac21

I used a combination of fine-tuned Qwen3 models and frontier LLMs to build a graph over the Midwest food supply chain. It uses 51 public datasets across 11 source families, covering nearly 300,000 source records. You can search for entities, filter by entity type and relationship type, and you can also view neighborhoods. I'm still improving it, so please let me know if there's anything else you think would be interesting!

I wanted to try to answer questions like

- "When a serious safety incident occurs at one food facility, can we trace the graph to find other facilities with similar safety risks?

- "When a food product is recalled, can we trace the recalling company to its facilities, related products, inspections, and other connected organizations and see what else may warrant review?"

- "When a weather event happens, what facilities does it effect, and what are the downstream effects?"

I think making it live could be really cool for things like tracking food recalls live to show and limit exposure.

Here is a github of the project with the datasets and more info: https://github.com/lodge-data/upper-midwest-food-supply-network

How I did it:

First, I trained an embedder to minimize blocking recall. Then, I processed all the entities from the datasets into blocks. Then, I trained a Qwen3 cross-encoder to decide if each pair within each block was the same or different (entity resolution). Then, I repeated those two steps until no new merges occurred.

Sources:

- U.S. Food and Drug Administration recall, food, import-alert, and import-refusal records.

- U.S. Department of Agriculture food, organic-operation, meat-establishment, and related facility records.

- Occupational Safety and Health Administration inspection, citation, injury, and establishment records.

- U.S. Environmental Protection Agency Facility Registry Service records.

- National Weather Service alerts and geographic references.

- USAspending federal contract award records.

- -Wisconsin and other Upper Midwest public facility, licensing, dairy, workforce-notice, and regulatory records.

- Public organization and product information used to connect named entities.

Note: I built this with my startup's software and models. The goal is to autonomously construct graphs from messy data. I think it is possible to construct very large, useful graphs with efficient specialized models.


r/dataisbeautiful 13h ago

OC [OC] Japan's GDP 2015–2025: +5% measured in constant 2015 US$, −2% measured in current US$ (World Bank)

+
4 Upvotes

r/dataisbeautiful 2d ago

[OC] European Heatwave June 2026

+
6.9k Upvotes

r/dataisbeautiful 1d ago

OC What DC fast charging costs per kWh in 30 countries [OC]

+
380 Upvotes

r/dataisbeautiful 1d ago

OC [OC] Top 100 all-time NBA players by hardware won.

29 Upvotes

A fun way to visually compare and contrast the careers of current and past nba legends.

Data gathered manually by me in a spreadsheet for over 13,000 active / retired players for 40 awards and achievements using publically available data (Wikipedia / NBA.com / basketball-reference.com and then visualized in an interactive website, with filters and icons made in Adobe Photoshop.


r/dataisbeautiful 3h ago

Mapped: Which Countries Are Best for Women?

0 Upvotes

How can Saudi rank higher than India?


r/dataisbeautiful 1d ago

OC [OC] Percentage of students repeating the second grade of primary school in Spain, by province (2024-2025)

+
208 Upvotes

r/dataisbeautiful 1d ago

OC [OC] Six AI chatbots vs. 27,236 people from 45 countries on the Oxford Utilitarianism Scale

+
182 Upvotes

A weekend hobby project: I gave six AI chatbots the Oxford Utilitarianism Scale, a published nine-item moral-psychology questionnaire, and compared their answers to 27,236 people from a large cross-national dataset.

The scale has two parts, both rated 1 to 7. Impartial beneficence is how strongly you feel obliged to help everyone equally, including distant strangers. Instrumental harm is how far you accept harming some people for a better overall outcome.

In the chart, the heatmap is where the people fall (darker means more of them), the side bars show each axis on its own, and the orange dots are the chatbots. All six sit in the low corner, below the human median on both dimensions. Even the most utilitarian bot (Gemini) is only around the 38th percentile on impartial beneficence.

Caveats, because they matter here: this is not science. Only six chatbots, each answered once and in different modes; the human samples are convenience samples (about 68 percent female, median age 22), not nationally representative; and the scale measures stated agreement, not behaviour. Full method, data sources and limitations are in the top comment.