TL;DR
New admin API endpoints that allow import of common data into a Fluxer instance, in support of migrating communities from other platforms like Discord to Fluxer.
LLM disclosure
All text written in this discussion including comments are my own. I use Claude Code to explore the codebase and do research, brainstorm ideas and explore alternative perspectives.
Motivation
A blocker I encountered for switching my Discord friend group over to Fluxer was that there is no way to migrate our history and memories. I've already created a third party tool that dumps all desired Discord data, but there's currently no clean way to import this dump into a Fluxer instance. The proposal is a new
/admin/import endpoint in
fluxer_api that adds API support for two use-cases:
- Support third-party automation of data import.
- Support for a hypothetical UI in the admin panel that allows nontechnical users to import message history into an instance.
The long-term vision this feature enables is to lower the friction involved in moving communities to Fluxer.
Any design ideas in this draft are framed from my currently limited point of view and are open to change as developers more experienced with the codebase weigh in.
What data types should the feature support?
Since I expect Discord to be the largest source of migrating communities, I'll focus on what users may want to bring over from there.
However, I think the implementation shouldn't assume anything about the data source.
I'll list Fluxer entities in what I perceive to be their natural order.
Some entities are cheap to implement in that they don't depend on other data being present:
- Emoji / Sticker
- Channel
- Role / Permissions
Any other entity will depend on their associated User being present. Since I personally don't know how identities will evolve in light of Federation, it's hard for me to design around this on my own, so I think this will require the most discussion.
I'll add my current thoughts on the matter in [Implementation Details](#implementation-details).
Assuming we've solved User seeding, we can move on to user-dependent entities:
- GuildMember
- => Link User-specific Roles
- Message
Some desired data types aren't supported yet in Fluxer:
- Poll message type
- Thread channel type
- Other?
Proposed changes
The operating assumption is that the instance has guild(s) pre-created by an admin before anything import-related happens.
I propose adding a new
ImportController in
fluxer_api that centralizes the discussed functionality and implements the necessary routes.
PUT /admin/import/emojis/:emoji_idPUT /admin/import/stickers/:sticker_idPUT /admin/import/channels/:channel_idPUT /admin/import/roles/:role_idPUT /admin/import/users/:user_idPUT /admin/import/members/:user_idPOST /admin/import/messages
Implementation details
Common
- I assume body carries
guild_id since all endpoints (except /users) are guild-specific. - Each entity's id is specified explicitly because it encodes the creation timestamp.
- I suggest reusing the existing
/channels/:channel_id/attachments for presigned URLs whenever binary assets are required (attachments, emojis, stickers, avatars). The added routes can then reference their related asset when a resource is added.
Security / impersonation
The endpoints sit behind admin API keys.
I see two different security situations as it pertains to impersonation, since
author_id is up to the caller:
1) An admin that has root access to the server the instance runs on. Technically any admin with this level of access to an instance already has direct database access so the concern isn't new.
2) An admin that isn't an infrastructure owner. The feature introduces a new security concern in this case.
Should imported content be distinguishable?
Conflict policy
To accommodate cases where a resource id already exists, I propose an
overwrite flag per-resource.
- False => the server denies if id matches existing resource
- True => server full replaces the resource
Caution is the responsibility of the caller. Existing admin API endpoints can be used to GET data and match existing content.
Messages
- The message endpoint is POST because batching makes most sense for hundreds of thousands of messages. A single message can just have batch size 1. Report an outcome per message as per the [Conflict Policy](#conflict-policy).
- I think imported messages should be search-indexed.
- Imported messages presumably shouldn't fire events to connected clients.
Users
My thoughts as someone who's just onboarding into the codebase. Let me know what I'm missing.
Fluxer has a concept of claimed/unclaimed users, derived from whether or not a user has a password. This would hypothetically allow us to seed unclaimed accounts into an instance. These accounts could then be claimed by the original user. Some questions arise:
- How do we make sure the right user claims the right account?
- I know there will be some account merging functionality in the future, so if a user already has a Fluxer identity somewhere else we probably want them to be able to merge with their seeded user to prevent them from having to handle 20 different accounts.
Notice that the proposed implementation allows messages to be linked to existing users in the instance, since the caller is free to assign user ids.
Discussion
I designed the feature against maintainability issues I encountered during the writing of my own import tool. I had used the repository layer (which worked, I successfully imported a full Discord server with 500k+ messages), but this wasn't durable long-term as I would like to make it public. I'm looking for general feedback on the shape of the feature to inform a first PR I'd like to implement. The main concerns I'd especially like feedback on are security and handling of users.
2 comments
Comment by @lucyrose39
Comment waiting for review
Comment by @ilianbronchart
TOS
There's precedent for my proposed feature (I actually didn't know the extent, these are useful references to have):- Zulip has importers for Mattermost, Slack, Rocket.Chat: See Here. It even specifically mentions importing data from a Rocket.Chat database dump.
- Mattermost has support for Bulk loading data
- Rocket.Chat has UI for Slack data imports and a CSV importer
I think how operators obtain their data, and the TOS contract between them and any platform are their own concern. Fluxer isn't pushing anyone toward breaking any platform's TOS since the proposed API is source-agnostic, which is a much softer implementation than some of the precedents above. I see data import as a fundamental feature any chat platform should have to give hosters more control over their software.Migration fidelity
My feature targets communities that value high migration fidelity. Say a server operator discovers Fluxer and is interested in the prospect of moving their community and the memories they built over the years. But, upon researching further, they learn that the only way to migrate data is lossy per-user dumps that likely leave historical user interactions as one-sided conversations. They are left with the decision to either start a fresh Fluxer instance and forget about history or look for alternatives that do allow data import. I personally don't think many admins/users will want to bother with the manual labor involved in moving a community on a per-user basis. By lossy I mean (in the case of Discord user data exports as an example):- No reactions
- Attachments are exported as expiring links (~24h), which are likely already lost by the time a user presents their data.
- Missing pins, embeds, rich content and other metadata
Side-note: Discord user data exports contain IPs and payment information, so we don't want to encourage people uploading that. Other data that is not covered by a per-user migration:Federation
I think these two features can work alongside each other. My import feature allows external data to flow into the Fluxer ecosystem in a general way. Once a user claims this data, they have the ability to do whatever they want with it. They can request their high-fidelity data to be moved to any federated instance they like or they can delete it if they so choose. In any case (whether for user uploaded data approved by an admin or for external admin tooling support), we need API endpoints to write these objects into an instance. I think my proposal addresses that.