Edit history

Data Import Feature Proposal has not been edited, so there are no earlier versions.

Current version | Original by Rex
Show

Data Import Feature Proposal

TL;DR

New admin API endpoints that allow import of common data into a Fluxer instance, in support of migrating communities from other platforms like Discord to Fluxer.

LLM disclosure

All text written in this discussion including comments are my own. I use Claude Code to explore the codebase and do research, brainstorm ideas and explore alternative perspectives.

Motivation

A blocker I encountered for switching my Discord friend group over to Fluxer was that there is no way to migrate our history and memories. I've already created a third party tool that dumps all desired Discord data, but there's currently no clean way to import this dump into a Fluxer instance. The proposal is a new /admin/import endpoint in fluxer_api that adds API support for two use-cases:
  1. Support third-party automation of data import.
  2. Support for a hypothetical UI in the admin panel that allows nontechnical users to import message history into an instance.
The long-term vision this feature enables is to lower the friction involved in moving communities to Fluxer. Any design ideas in this draft are framed from my currently limited point of view and are open to change as developers more experienced with the codebase weigh in.

What data types should the feature support?

Since I expect Discord to be the largest source of migrating communities, I'll focus on what users may want to bring over from there. However, I think the implementation shouldn't assume anything about the data source. I'll list Fluxer entities in what I perceive to be their natural order. Some entities are cheap to implement in that they don't depend on other data being present:
  • Emoji / Sticker
  • Channel
  • Role / Permissions
Any other entity will depend on their associated User being present. Since I personally don't know how identities will evolve in light of Federation, it's hard for me to design around this on my own, so I think this will require the most discussion. I'll add my current thoughts on the matter in [Implementation Details](#implementation-details). Assuming we've solved User seeding, we can move on to user-dependent entities:
  • GuildMember
    • => Link User-specific Roles
  • Message
    • Attachment
    • Embed
    • Reaction
Some desired data types aren't supported yet in Fluxer:
  • Poll message type
  • Thread channel type
  • Other?

Proposed changes

The operating assumption is that the instance has guild(s) pre-created by an admin before anything import-related happens. I propose adding a new ImportController in fluxer_api that centralizes the discussed functionality and implements the necessary routes.
  • PUT /admin/import/emojis/:emoji_id
  • PUT /admin/import/stickers/:sticker_id
  • PUT /admin/import/channels/:channel_id
  • PUT /admin/import/roles/:role_id
  • PUT /admin/import/users/:user_id
  • PUT /admin/import/members/:user_id
  • POST /admin/import/messages

Implementation details

Common
  • I assume body carries guild_id since all endpoints (except /users) are guild-specific.
  • Each entity's id is specified explicitly because it encodes the creation timestamp.
  • I suggest reusing the existing /channels/:channel_id/attachments for presigned URLs whenever binary assets are required (attachments, emojis, stickers, avatars). The added routes can then reference their related asset when a resource is added.
Security / impersonation
The endpoints sit behind admin API keys. I see two different security situations as it pertains to impersonation, since author_id is up to the caller: 1) An admin that has root access to the server the instance runs on. Technically any admin with this level of access to an instance already has direct database access so the concern isn't new. 2) An admin that isn't an infrastructure owner. The feature introduces a new security concern in this case. Should imported content be distinguishable?
Conflict policy
To accommodate cases where a resource id already exists, I propose an overwrite flag per-resource.
  • False => the server denies if id matches existing resource
  • True => server full replaces the resource
Caution is the responsibility of the caller. Existing admin API endpoints can be used to GET data and match existing content.
Messages
  • The message endpoint is POST because batching makes most sense for hundreds of thousands of messages. A single message can just have batch size 1. Report an outcome per message as per the [Conflict Policy](#conflict-policy).
  • I think imported messages should be search-indexed.
  • Imported messages presumably shouldn't fire events to connected clients.
Users
My thoughts as someone who's just onboarding into the codebase. Let me know what I'm missing. Fluxer has a concept of claimed/unclaimed users, derived from whether or not a user has a password. This would hypothetically allow us to seed unclaimed accounts into an instance. These accounts could then be claimed by the original user. Some questions arise:
  • How do we make sure the right user claims the right account?
  • I know there will be some account merging functionality in the future, so if a user already has a Fluxer identity somewhere else we probably want them to be able to merge with their seeded user to prevent them from having to handle 20 different accounts.
Notice that the proposed implementation allows messages to be linked to existing users in the instance, since the caller is free to assign user ids.

Discussion

I designed the feature against maintainability issues I encountered during the writing of my own import tool. I had used the repository layer (which worked, I successfully imported a full Discord server with 500k+ messages), but this wasn't durable long-term as I would like to make it public. I'm looking for general feedback on the shape of the feature to inform a first PR I'd like to implement. The main concerns I'd especially like feedback on are security and handling of users.