Data and AI Community
A data and machine learning community: datasets, model discussion, notebook help, a papers forum and careers.
- 8 categories
- 35 channels
- 15 roles
- Medium
Choose a server, review what will be created or reused, then confirm. Merge keeps what you have; replace rebuilds the server. Vetox takes a safety backup first either way.
Server structure
Channels
ENTRY
- community-guide
Which category is for what, and how the paper club works.
- house-rules
Cite sources, state licences, no recruiting outside #data-jobs.
- field-roles
Data Engineering, Machine Learning, Analytics, Research, and the ping roles.
- community-bulletin
Sessions, guests and changes to the server.
- data-lounge
Everything that is not a question, a paper or a job.
DATASETS
- dataset-share
One thread per dataset: source, licence, size, format, what it is good for.
- data-cleaning-help
Missing values, bad encodings, duplicate rows and the join that multiplied everything.
- data-sources-and-apis
Where to get data, and how the API behaves once you do.
- licensing-and-ethics
Can you use it, can you publish it, should you.
MODELS
- model-discussion
Architectures, approaches and the arguments that never end.
- training-and-fine-tuning
One thread per problem: data size, setup, what the loss is doing.
- benchmarks-and-evals
How to measure it, and why the number you got is suspicious.
- deployment-and-serving
Getting a model into production and keeping it there.
- prompting-and-agents
Prompt design, tool use and agent loops. Share what worked and what looped forever.
NOTEBOOKS AND CODE
- notebook-help
The stuck cell. Paste the code and the traceback, tag the topic.
- pipeline-review
Post a pipeline or a script for a second pair of eyes.
- sql-corner
Queries, plans and window functions.
- visualisation-critique
Post the chart, say what it should show, get told what it shows.
PAPERS
- paper-discussion
One thread per paper. Link it, summarise it in three lines, then argue.
- paper-club-schedule
The next paper, the date and the stage link. Mentions Paper Club Pings.
- paper-feed
Vetox posts new preprints from the feeds members follow.
- reproductions
Trying to make a published result hold. Post what you got and how far off it was.
CAREER
- data-jobs
One thread per role: company, location, range, link. Mark it filled.
- interview-prep
Take-homes, case studies and the SQL round.
- portfolio-reviews
Projects and profiles. Say which role you are aiming at.
- career-questions
Titles, moves, salaries and whether the degree matters.
VOICE
- Data Hangout
- Paper Club
- Study Session
- Hack Session
- AFK
STAFF
- staff-lounge
Reports, role grants and the recruiter of the week.
- moderation-queue
One message per report or removed dataset: link, reason, outcome.
- community-logsHidden from @everyone
Vetox posts moderation and join logs here.
- Staff Room
Roles
- OrganiserAdministrator
- AdminAdministrator
- ModeratorModerator
- Paper Club HostModerator
- Researcher
- Practitioner
- Student
- Data Engineering
- Machine Learning
- Analytics
- Research
- Job Pings
- Paper Club Pings
- Muted
Overview
A server for people who work with data for a living or want to: analysts, data engineers, machine learning practitioners, researchers and students. It separates the four conversations that usually collapse into one. Datasets holds a #dataset-share forum tagged by data type, #data-cleaning-help for the unglamorous part, #data-sources-and-apis and #licensing-and-ethics, because where data came from matters as much as what is in it. Models has #model-discussion for the open-ended arguments, a #training-and-fine-tuning forum tagged Training, Fine-tuning, Evaluation and Inference with a Solved tag, #benchmarks-and-evals, #deployment-and-serving and #prompting-and-agents. Notebooks and Code is where the actual work gets unstuck: #notebook-help is a forum tagged Pandas, Visualisation, SQL and Pipelines, with #pipeline-review, #sql-corner and #visualisation-critique beside it. Papers has a #paper-discussion forum, a #paper-club-schedule that only the club host posts to, a #paper-feed and #reproductions for people trying to get a result to hold. Career keeps #data-jobs as a forum so postings stay findable, alongside #interview-prep and #portfolio-reviews. Voice has a Paper Club stage and study and hack rooms.
When to use it
Your data community has one channel where a question about a broken pipeline, a link to a new preprint and a job posting arrive within the same minute, and the people who could answer each have muted it. Splitting datasets, models, notebooks, papers and careers into their own categories means each conversation has an audience that chose it, the paper club has a schedule and a stage, and the code questions live in forums where a solved thread is worth more than a scrolled-away answer.
What makes it different
- #dataset-share is a forum tagged Tabular, Text, Images, Time series, Audio and Synthetic, with licensing beside it.
- #training-and-fine-tuning is a forum tagged Training, Fine-tuning, Evaluation and Inference, with a Solved tag.
- #notebook-help is a forum tagged Pandas, Visualisation, SQL and Pipelines, so the stuck cell finds the right reader.
- Paper Club has a forum, a host-only schedule channel, a read-only feed and a stage for the weekly session.
- #data-jobs is a forum, so a posting stays findable for weeks instead of scrolling away in a day.
Recommended Vetox setup
Self-Roles
Post Self-Roles in #field-roles for Data Engineering, Machine Learning, Analytics and Research, plus Paper Club Pings and Job Pings, so the stage session and new postings reach the people who asked.
Reminders
Schedule a Reminder in #paper-club-schedule two days before each session with the paper link, and one an hour before it in #paper-discussion.
Notifications
Use Notifications with the RSS feeds of the preprint categories your members follow, posted to #paper-feed, so #paper-discussion stays for discussion.
Auto-Mod
Enable Auto-Mod in #data-jobs and #career-questions against link spam; recruiter bots are the main moderation load in a data server.
Questions
Why is there no channel per framework or library?
Because data questions are about the problem, not the tool: a pipeline that drops rows is the same problem in any framework. The forums are tagged by task - Training, Evaluation, SQL, Pipelines - and the tag tells the reader what kind of help is needed. Add a framework channel only when one tool dominates your members' work.
How does the paper club run?
The host posts the next paper in #paper-club-schedule, which is read-only for everyone else, and opens a thread for it in #paper-discussion. Discussion happens in the thread during the week and on the Paper Club stage at the session. #reproductions is for the people who go further and try to make the result hold.
Can people post datasets they scraped?
Only with the licence and source stated, which is why #licensing-and-ethics sits next to #dataset-share and why the share forum asks for both in the first post. Moderators remove a dataset whose provenance cannot be explained.
Common mistakes
- Treating the server as a news feed, where every new preprint is posted into the same channel people ask questions in.
- Letting job postings into general chat, which brings recruiters who never leave and buries the questions.
Consider instead
- Programming Community Your members are general software developers and data is one topic among many rather than the reason the server exists.
More in Developers
View allDev Team Internal
FeaturedA private server for one engineering team: async standups, sprint rituals, code review, incidents and deploy feeds.
DevOps and Cloud Engineers
FeaturedDevOps and cloud engineers: provider channels, Kubernetes help, CI/CD, incident war stories and certifications.
Open-Source Project Hub
FeaturedA home for one open-source project: contributor onboarding, issue triage, RFCs, release notes and read-only CI feeds.
Programming Community
FeaturedA public programming server: per-language channels, a tagged help forum, a jobs board and a project showcase.
0.0k
Servers
0+
Commands
0
Users
0
Languages
