Data and AI Community

A data and machine learning community: datasets, model discussion, notebook help, a papers forum and careers.

  • 8 categories
  • 35 channels
  • 15 roles
  • Medium
Download Discord JSON

Choose a server, review what will be created or reused, then confirm. Merge keeps what you have; replace rebuilds the server. Vetox takes a safety backup first either way.

No uses yetNo views yet

Server structure

Channels

ENTRY

  • community-guide

    Which category is for what, and how the paper club works.

  • house-rules

    Cite sources, state licences, no recruiting outside #data-jobs.

  • field-roles

    Data Engineering, Machine Learning, Analytics, Research, and the ping roles.

  • community-bulletin

    Sessions, guests and changes to the server.

  • data-lounge

    Everything that is not a question, a paper or a job.

DATASETS

  • dataset-share

    One thread per dataset: source, licence, size, format, what it is good for.

  • data-cleaning-help

    Missing values, bad encodings, duplicate rows and the join that multiplied everything.

  • data-sources-and-apis

    Where to get data, and how the API behaves once you do.

  • licensing-and-ethics

    Can you use it, can you publish it, should you.

MODELS

  • model-discussion

    Architectures, approaches and the arguments that never end.

  • training-and-fine-tuning

    One thread per problem: data size, setup, what the loss is doing.

  • benchmarks-and-evals

    How to measure it, and why the number you got is suspicious.

  • deployment-and-serving

    Getting a model into production and keeping it there.

  • prompting-and-agents

    Prompt design, tool use and agent loops. Share what worked and what looped forever.

NOTEBOOKS AND CODE

  • notebook-help

    The stuck cell. Paste the code and the traceback, tag the topic.

  • pipeline-review

    Post a pipeline or a script for a second pair of eyes.

  • sql-corner

    Queries, plans and window functions.

  • visualisation-critique

    Post the chart, say what it should show, get told what it shows.

PAPERS

  • paper-discussion

    One thread per paper. Link it, summarise it in three lines, then argue.

  • paper-club-schedule

    The next paper, the date and the stage link. Mentions Paper Club Pings.

  • paper-feed

    Vetox posts new preprints from the feeds members follow.

  • reproductions

    Trying to make a published result hold. Post what you got and how far off it was.

CAREER

  • data-jobs

    One thread per role: company, location, range, link. Mark it filled.

  • interview-prep

    Take-homes, case studies and the SQL round.

  • portfolio-reviews

    Projects and profiles. Say which role you are aiming at.

  • career-questions

    Titles, moves, salaries and whether the degree matters.

VOICE

  • Data Hangout
  • Paper Club
  • Study Session
  • Hack Session
  • AFK

STAFF

  • staff-lounge

    Reports, role grants and the recruiter of the week.

  • moderation-queue

    One message per report or removed dataset: link, reason, outcome.

  • community-logsHidden from @everyone

    Vetox posts moderation and join logs here.

  • Staff Room

Roles

  • OrganiserAdministrator
  • AdminAdministrator
  • ModeratorModerator
  • Paper Club HostModerator
  • Researcher
  • Practitioner
  • Student
  • Data Engineering
  • Machine Learning
  • Analytics
  • Research
  • Job Pings
  • Paper Club Pings
  • Muted

Overview

A server for people who work with data for a living or want to: analysts, data engineers, machine learning practitioners, researchers and students. It separates the four conversations that usually collapse into one. Datasets holds a #dataset-share forum tagged by data type, #data-cleaning-help for the unglamorous part, #data-sources-and-apis and #licensing-and-ethics, because where data came from matters as much as what is in it. Models has #model-discussion for the open-ended arguments, a #training-and-fine-tuning forum tagged Training, Fine-tuning, Evaluation and Inference with a Solved tag, #benchmarks-and-evals, #deployment-and-serving and #prompting-and-agents. Notebooks and Code is where the actual work gets unstuck: #notebook-help is a forum tagged Pandas, Visualisation, SQL and Pipelines, with #pipeline-review, #sql-corner and #visualisation-critique beside it. Papers has a #paper-discussion forum, a #paper-club-schedule that only the club host posts to, a #paper-feed and #reproductions for people trying to get a result to hold. Career keeps #data-jobs as a forum so postings stay findable, alongside #interview-prep and #portfolio-reviews. Voice has a Paper Club stage and study and hack rooms.

When to use it

Your data community has one channel where a question about a broken pipeline, a link to a new preprint and a job posting arrive within the same minute, and the people who could answer each have muted it. Splitting datasets, models, notebooks, papers and careers into their own categories means each conversation has an audience that chose it, the paper club has a schedule and a stage, and the code questions live in forums where a solved thread is worth more than a scrolled-away answer.

What makes it different

  • #dataset-share is a forum tagged Tabular, Text, Images, Time series, Audio and Synthetic, with licensing beside it.
  • #training-and-fine-tuning is a forum tagged Training, Fine-tuning, Evaluation and Inference, with a Solved tag.
  • #notebook-help is a forum tagged Pandas, Visualisation, SQL and Pipelines, so the stuck cell finds the right reader.
  • Paper Club has a forum, a host-only schedule channel, a read-only feed and a stage for the weekly session.
  • #data-jobs is a forum, so a posting stays findable for weeks instead of scrolling away in a day.

Recommended Vetox setup

  • Self-Roles

    Post Self-Roles in #field-roles for Data Engineering, Machine Learning, Analytics and Research, plus Paper Club Pings and Job Pings, so the stage session and new postings reach the people who asked.

  • Reminders

    Schedule a Reminder in #paper-club-schedule two days before each session with the paper link, and one an hour before it in #paper-discussion.

  • Notifications

    Use Notifications with the RSS feeds of the preprint categories your members follow, posted to #paper-feed, so #paper-discussion stays for discussion.

  • Auto-Mod

    Enable Auto-Mod in #data-jobs and #career-questions against link spam; recruiter bots are the main moderation load in a data server.

Questions

Why is there no channel per framework or library?

Because data questions are about the problem, not the tool: a pipeline that drops rows is the same problem in any framework. The forums are tagged by task - Training, Evaluation, SQL, Pipelines - and the tag tells the reader what kind of help is needed. Add a framework channel only when one tool dominates your members' work.

How does the paper club run?

The host posts the next paper in #paper-club-schedule, which is read-only for everyone else, and opens a thread for it in #paper-discussion. Discussion happens in the thread during the week and on the Paper Club stage at the session. #reproductions is for the people who go further and try to make the result hold.

Can people post datasets they scraped?

Only with the licence and source stated, which is why #licensing-and-ethics sits next to #dataset-share and why the share forum asks for both in the first post. Moderators remove a dataset whose provenance cannot be explained.

Common mistakes

  • Treating the server as a news feed, where every new preprint is posted into the same channel people ask questions in.
  • Letting job postings into general chat, which brings recruiters who never leave and buries the questions.

Consider instead

  • Programming Community Your members are general software developers and data is one topic among many rather than the reason the server exists.

Do you have any doubts?

Our team is there for you

0.0k

Servers

0+

Commands

0

Users

0

Languages