Artificial intelligence enhanced mobile camera interface
Techniques for artificial intelligence (AI) enhanced mobile camera interfaces are described and are implementable to save users from providing complex and tedious inputs to take pictures. In implementations, a mobile device modifies, based at least in part on user instructions to the mobile device and using an artificial intelligence model, an image preview of a camera field of view presented on the mobile device. The image preview is replaced with a modified image preview generated by the artificial intelligence model based on the user instructions. The mobile device generates, responsive to detecting a capture command and using the artificial intelligence model, a captured image for display as a final image within a user interface based at least in part on the modified image preview.
1 . A mobile device, comprising:
at least one memory; and
at least one processor coupled with the at least one memory and configured to cause the mobile device to:
modify, based at least in part on user instructions to the mobile device and using an artificial intelligence model, an image preview of a camera field of view presented on the mobile device;
replace the image preview with a modified image preview generated by the artificial intelligence model based on the user instructions; and
generate, responsive to detecting a capture command and using the artificial intelligence model, a captured image for display as a final image within a user interface based at least in part on the modified image preview.
2 . The mobile device of claim 1 , wherein the user instructions are extracted by the artificial intelligence model from a voice command received at the user interface.
3 . The mobile device of claim 2 , wherein the voice command indicates one or more instructions for modifying the image preview, wherein the artificial intelligence model is configured to iteratively process the one or more instructions to replace the image preview with a different image generated by the artificial intelligence model for each different instruction, and wherein the different image generated for a final instruction corresponds to the modified image.
4 . The mobile device of claim 1 , wherein the artificial intelligence model is configured to generate the modified image by at least one of adding, removing, or manipulating image features requested by the user instructions.
5 . The mobile device of claim 1 , wherein the artificial intelligence model is configured to generate the modified image by applying image enhancements requested by the user instructions.
6 . The mobile device of claim 1 , wherein the artificial intelligence model is configured to generate the modified image based further on semantic information extracted from the image preview.
7 . The mobile device of claim 1 , wherein the artificial intelligence model is configured to generate the modified image based further on a textual prompt combining a transcription of the user instructions with semantic information extracted from the image preview.
8 . The mobile device of claim 1 , wherein the artificial intelligence model is configured to generate the modified image based further on a conversation history of the user instructions received at the user interface during past user interactions.
9 . The mobile device of claim 1 , wherein the artificial intelligence model is configured to generate the modified image based further on contextual information associated with the mobile device and user preference information.
10 . The mobile device of claim 9 , wherein the contextual information is inferred from portions of an environment shown in the image preview.
11 . A system comprising:
at least one memory; and
at least one processor coupled with the at least one memory and configured to cause the system to:
modify, based at least in part on user instructions to the system and using an artificial intelligence model, an image preview of a camera field of view presented by the system;
replace the image preview with a modified image preview generated by the artificial intelligence model based on the user instructions; and
generate, responsive to detecting a capture command and using the artificial intelligence model, a captured image for display as a final image within a user interface based at least in part on the modified image preview.
12 . The system of claim 11 , wherein the artificial intelligence model is configured to generate the modified image by at least one of adding, removing, or manipulating image features requested by the user instructions, and by applying image enhancements requested by the user instructions.
13 . The system of claim 11 , wherein the artificial intelligence model is configured to generate the modified image based further on semantic information extracted from the image preview.
14 . The system of claim 11 , wherein the artificial intelligence model is configured to generate the modified image based further on a textual prompt combining a transcription of the user instructions with semantic information extracted from the image preview.
15 . The system of claim 11 , wherein the artificial intelligence model is configured to generate the modified image based further on a conversation history of user instructions received at the user interface during past user interactions.
16 . The system of claim 11 , wherein the artificial intelligence model is configured to generate the modified image based further on contextual information inferred from portions of the image preview and user preference information.
17 . A method performed by a mobile device, the method comprising:
executing, by at least one processor of a mobile device, a camera application including an artificial intelligence model that processes voice commands received at a user interface for controlling a camera of the mobile device;
presenting, by the at least one processor, a camera viewfinder within the user interface including an image preview of a field of view of the camera;
detecting, by the at least one processor, a voice command at the user interface that indicates user instructions for modifying the image preview;
replacing, by the at least one processor, the image preview with a modified image generated by the artificial intelligence model based on the user instructions;
receiving, by the at least one processor, a capture command at the user interface that causes the camera to output a captured image of the field of view;
using, by the at least one processor, the artificial intelligence model to generate a final image by applying same modifications to the captured image as applied by the artificial intelligence model to the image preview to generate the modified image; and
storing, by the at least one processor, the final image within a memory.
18 . The method of claim 17 , wherein the voice command indicates a plurality of different instructions for modifying the image preview, and the replacing includes iteratively processing the different instructions to replace the image preview with a different image generated by the artificial intelligence model for each different instruction.
19 . The method of claim 18 , wherein the voice command indicates a specific order to the different instructions, and wherein the iteratively processing includes iteratively processing the different instructions in that specific order.
20 . The method of claim 17 , wherein the replacing includes using the artificial intelligence model for at least one of adding image features, removing the image features, manipulating the image features, or applying image enhancements.