Search labeling system for enhanced content indexing and retrieval
A system creates a block within a digital workspace. The block includes text that is assigned a text instance identifier (ID). The system can store a single conflict-free replicated data type (CRDT) text slice tree for the block, where each node represents a text slice corresponding to a segment of the text. A database table can store the text instance ID and a search label for each text slice. The system can, in response to a text editing operation, segment the text into two contiguous text slices, represented as first and second child nodes of the text slice tree, and assign a child search label that comprises a parent search label appended with different characters. A combination of the text instance ID and a search label uniquely identifies a text slice in the digital workspace.
1 . A non-transitory, computer-readable storage medium comprising instructions recorded thereon, wherein the instructions, when executed by at least one data processor of a system, cause the system to:
create a block within a digital workspace that includes multiple blocks,
wherein the block is a discrete unit of content including text that is a conflict-free replicated data type (CRDT) instance that is assigned a text instance identifier (ID);
store a single text slice tree in a database record linked to the block,
wherein the text slice tree has a hierarchical structure in which each node represents a text slice corresponding to a segment of the text, and
wherein a database table stores the text instance ID for the text and a search label for each text slice including a parent search label for a root node, each search label including a string identifier;
in response to a text editing operation to split the text of the block:
segment the text into two contiguous text slices, represented as a first child node and a second child node in the text slice tree;
propagate the parent search label to both the first and second child nodes;
assign to the first child node a first child search label that comprises the parent search label as a prefix appended with a first character;
assign to the second child node a second child search label that comprises the parent search label as a prefix appended with a second character, the second character being different from the first character; and
store the first and second child search labels in the database table,
wherein a combination of the text instance ID and a particular search label for a particular node of the text slice tree uniquely identifies a corresponding particular text slice in the digital workspace, and
wherein associations between search labels and their respective text slices are preserved throughout text editing operations performed on the text of the block.
2 . The non-transitory, computer-readable storage medium of claim 1 , wherein the text editing operation is a first text editing operation, and wherein the system is further caused to:
initiate a second text editing operation on a target location of the text of the block; and
in response to the second text editing operation:
identify the text instance ID and a target search label of a target segment;
query the database table for any text slices matching the text instance ID and a target prefix of the target search label,
wherein a set of matching text slices are retrieved for matching the target prefix and the target search label;
perform a search on the set of matching text slices to identify a particular text slice having a particular search label that contains the target prefix and a string identifier for the target search label; and
apply the text editing operation to the particular text slice.
3 . The non-transitory, computer-readable storage medium of claim 1 , wherein each search label is encoded as a text string comprising printable ASCII characters using a bit-packing scheme that encodes multiple split operations performed on text per byte.
4 . The non-transitory, computer-readable storage medium of claim 1 , wherein the system is further caused to, in response to the text editing operation:
insert a split character at a position of a parent text slice where the parent text slice is to be split;
designate a sequence of characters from a start point of the parent text slice to the split character as a first child slice; and
designate a sequence of characters from the split character to an end point of the parent text slice as a second child slice.
5 . The non-transitory, computer-readable storage medium of claim 1 , wherein the block is a first block, the text slice tree is a first text slice tree, and wherein the system is further caused to:
merge the first block with a second block by moving text slices from the first block to the second block while retaining respective search labels and text instance IDs of the moved text slices; and
reparent the first text slice tree of the first block as a child node of a last text slice in a second text slice tree of the second block.
6 . The non-transitory, computer-readable storage medium of claim 1 , wherein the system is further caused to:
synchronize multiple text editing operations performed asynchronously, by distributed client instances, on the text of the block by using the text instance ID and search labels to search the database table to enable asynchronous collaborative editing without data loss.
7 . The non-transitory, computer-readable storage medium of claim 1 , wherein the system is further caused to:
while a client of the digital workspace is offline:
store one or more text editing operations performed at that client;
assign a provisional identifier to the one or more text editing operations as offline operations; and
once the client of the digital workspace goes online after being offline:
synchronize the offline operations by using the provisional identifier and search labels to enable resolution of concurrent edits made by multiple clients while the client was offline.
8 . The non-transitory, computer-readable storage medium of claim 1 , wherein the system is further caused to:
store each search label of each text slice in the database table as metadata for the text slice,
wherein the metadata includes format, attribution, originating block, version information, operational history, and a unique CRDT instance identifier.
9 . The non-transitory, computer-readable storage medium of claim 1 , wherein the system is further caused to:
determine an order of a set of text slices within the block based on respective search labels associated with the set of text slices; and
render the block by displaying the text slices in accordance with the order.
10 . A method comprising:
creating a block within a digital workspace that includes multiple blocks,
wherein the block is a discrete unit of content including text that is a conflict-free replicated data type (CRDT) instance that is assigned a text instance identifier (ID);
storing a single text slice tree in a database record linked to the block,
wherein the text slice tree has a hierarchical structure in which each node represents a text slice corresponding to a segment of the text, and
wherein a database table stores the text instance ID for the text and a search label for each text slice including a parent search label for a root node, each search label including a string identifier;
in response to a text editing operation to split the text of the block:
segmenting the text into two contiguous text slices, represented as a first child node and a second child node in the text slice tree;
propagating the parent search label to both the first and second child nodes;
assigning to the first child node a first child search label that comprises the parent search label as a prefix appended with a first character;
assigning to the second child node a second child search label that comprises the parent search label as a prefix appended with a second character, the second character being different from the first character; and
storing the first and second child search labels in the database table,
wherein a combination of the text instance ID and a particular search label for a particular node of the text slice tree uniquely identifies a corresponding particular text slice in the digital workspace, and
wherein associations between search labels and their respective text slices are preserved throughout text editing operations performed on the text of the block.
11 . The method of claim 10 , further comprising:
initiating a second text editing operation on a target location of the text of the block; and
in response to the second text editing operation:
identifying the text instance ID and a target search label of a target segment;
querying the database table for any text slices matching the text instance ID and a target prefix of the target search label,
wherein a set of matching text slices are retrieved for matching the target prefix and the target search label;
performing a search on the set of matching text slices to identify a particular text slice having a particular search label that contains the target prefix and a string identifier for the target search label; and
applying the text editing operation to the particular text slice.
12 . The method of claim 10 , wherein each search label is encoded as a text string comprising printable ASCII characters using a bit-packing scheme that encodes multiple split operations performed on text per byte.
13 . The method of claim 10 , further comprising:
inserting a split character at a position of a parent text slice where the parent text slice is to be split;
designating a sequence of characters from a start point of the parent text slice to the split character as a first child slice; and
designating a sequence of characters from the split character to an end point of the parent text slice as a second child slice.
14 . The method of claim 10 , wherein the block is a first block, the text slice tree is a first text slice tree, and the method further comprises:
merging the first block with a second block by moving text slices from the first block to the second block while retaining respective search labels and text instance IDs of the moved text slices; and
reparenting the first text slice tree of the first block as a child node of a last text slice in a second text slice tree of the second block.
15 . The method of claim 10 , further comprising:
synchronizing multiple text editing operations performed asynchronously, by distributed client instances, on the text of the block by using the text instance ID and search labels to search the database table to enable asynchronous collaborative editing without data loss.
16 . A system comprising:
at least one hardware processor; and
at least one non-transitory memory storing instructions, which, when executed by the at least one hardware processor, cause the system to:
create a block within a digital workspace that includes multiple blocks,
wherein the block is a discrete unit of content including text that is a conflict-free replicated data type (CRDT) instance that is assigned a text instance identifier (ID);
store a single text slice tree in a database record linked to the block,
wherein the text slice tree has a hierarchical structure in which each node represents a text slice corresponding to a segment of the text, and
wherein a database table stores the text instance ID for the text and a search label for each text slice including a parent search label for a root node, each search label including a string identifier;
in response to a text editing operation to split the text of the block:
segment the text into two contiguous text slices, represented as a first child node and a second child node in the text slice tree;
propagate the parent search label to both the first and second child nodes;
assign to the first child node a first child search label that comprises the parent search label as a prefix appended with a first character;
assign to the second child node a second child search label that comprises the parent search label as a prefix appended with a second character, the second character being different from the first character; and
store the first and second child search labels in the database table,
wherein a combination of the text instance ID and a particular search label for a particular node of the text slice tree uniquely identifies a corresponding particular text slice in the digital workspace, and
wherein associations between search labels and their respective text slices are preserved throughout text editing operations performed on the text of the block.
17 . The system of claim 16 , wherein the system is further caused to:
initiate a second text editing operation on a target location of the text of the block; and
in response to the second text editing operation:
identify the text instance ID and a target search label of a target segment;
query the database table for any text slices matching the text instance ID and a target prefix of the target search label,
wherein a set of matching text slices are retrieved for matching the target prefix and the target search label;
perform a search on the set of matching text slices to identify a particular text slice having a particular search label that contains the target prefix and a string identifier for the target search label; and
apply the text editing operation to the particular text slice.
18 . The system of claim 16 , wherein each search label is encoded as a text string comprising printable ASCII characters using a bit-packing scheme that encodes multiple split operations performed on text per byte.
19 . The system of claim 16 , wherein the system is further caused to, in response to the text editing operation:
insert a split character at a position of a parent text slice where the parent text slice is to be split;
designate a sequence of characters from a start point of the parent text slice to the split character as a first child slice; and
designate a sequence of characters from the split character to an end point of the parent text slice as a second child slice.